Stop Asking LLMs for Booleans: Why Your Code Needs Jev

Utkarsh Tiwari
8 min read
jevllmaiengineeringcoding

Over the last few days, one word keeps showing up in my feed: Jev.

I spent some time reading and exploring about it and, Here's what I gathered, in plain words.

If you've built anything with LLMs, you know this pain.

You need a simple decision. You write: "Is this user angry? Answer ONLY Yes or No."

The model replies: "Based on the text provided, the user seems frustrated, so I would say… Yes."

Now you're writing regex, forcing JSON and adding retries. All for one boolean.

Jev is built for exactly that problem.

chatty robot vs jev

What Jev Actually Is

In plain developer terms, Jev is a generalized classifier, not a text generator.

Most AI models are built to talk. Jev is built to act like a smart, probabilistic if/else statement. Instead of generating a paragraph you have to parse with regex, it accepts unstructured text or JSON (the "state") and immediately maps it to predefined schemas—like category labels, ratings, or booleans—in a single, high-speed pass.

The vision behind it:

TypeSafe founder Diogo Almeida built Jev around Daniel Kahneman's concept of System 1 vs. System 2 thinking:

  • System 2 (LLMs): Slow, step-by-step reasoning. Great for drafting emails, debugging complex code, or writing analysis.
  • System 1 (Jev): Instant, automatic reflex. Perfect for deciding "Is this spam?", "Which team handles this ticket?", or "Which tool should the agent run next?"

system1 vs system2

Rather than training the model to sound polite or conversational using standard RLHF, TypeSafe trained Jev using RLCD (Reinforcement Learning for Calibrated Decisions). The goal isn't to sound convincing—it's to give your code a fast, deterministic answer with a confidence score you can safely rely on.

The Three Building Blocks

Jev only answers in three formats:

  • Choice: pick one option from a list you define. It returns a probability distribution across every option (essentially a typed switch statement).
  • Score: rate something on an ordered scale you define.
  • Noul: a yes/no question, answered as a probability between 0 and 1.

Because you declare the possible answers up front, the output is always valid. There is no broken JSON, no schema drift, and no surprise paragraph.

Why People Are Excited

Here's what TypeSafe reports on its site:

  • 193.6x faster and 444.6x cheaper than LLMs on System One style workflows. Their own benchmark example shows $0.000081 in 0.114s vs $0.01388 in 8.566s.
  • $42 per billion input tokens (~4 rupees per 1M tokens). Output tokens are 100% free because it doesn't generate text.
  • 70–500 ms typical response times.
  • Parallel multi-question evaluation: you can pass several questions at once across the same state, and Jev evaluates them concurrently in a single pass.
  • Calibrated confidence scores: your code can set safe automated thresholds, like "auto-act above 0.95, otherwise route to a human review queue."
  • Zero hallucinations by design: because the model is bound to your schema, it can't invent facts or wander off script. Even when it is uncertain, that uncertainty shows up directly in the confidence score.

(Note: These figures come from TypeSafe's own site and launch materials. Take them as promising, not independently audited benchmarks.)

latency and cost vs llm

From "Feature" to "Primitive"

Right now, most teams treat AI as a heavy, standalone feature—a chatbot window, a copilot drawer, or a slow background worker.

Jev pushes AI toward becoming a basic programming primitive, just like an if/else statement or a database query.

Instead of writing brittle regex rules or paying for a full LLM call, you can write:

if jev("is_this_customer_angry", state) > 0.9: escalate_to_manager()

This bridges the gap between messy, unstructured real-world data (raw customer emails, chaotic Discord chats, multi-gigabyte server logs) and rigid, structured code logic.

What People Built in the First Week

A community repository called awesome-jev-use-cases tracked 74 demos in the first five days alone. A few practical setups that stood out:

  • AI Agent reflex routers: in complex agentic workflows, LLMs spend huge amounts of time and money just answering "Which tool should I call next?" or "Did this subtask finish?" Jev handles tool selection and state checks in 100 ms.
  • Claude Code context compaction: a plugin that scores tool calls and purges stale context before hitting the token ceiling.
  • Browser automation: a Browser Use project where Jev inspects page state and picks the next UI click or keystroke.
  • Email triage: one demo sorted 500 messy support emails for about 3.5 cents.
  • Real-time content moderation & slop detection: scoring social feeds, Twitch streams, and live chat for spam or toxicity before the UI even renders.
  • PostgreSQL jev() extension: querying and filtering unstructured table rows using plain-language classification.
  • Micro-decision games: powering quick state loops for retro game bots (Doom, Tetris, Subway Surfers). awesome-jev-projects

A Simple Example: Support Ticket Triage

Say your team receives hundreds of incoming tickets a day, and you want them sorted automatically without writing fragile rule engines.

You send Jev the ticket text as the "state" and ask three targeted questions:

  • A Choice: is this billing, bug report, sales, or other?
  • A Noul: does this need a response within 2 hours?
  • A Score: how frustrated is the user (calm, annoyed, angry)?

Jev returns typed values and confidence scores. All workflow decisions stay in regular, testable code:

  • If the Noul is above 0.9 and the Choice is billing, route directly to finance as high-priority.
  • If confidence on the Choice is under 0.6, push it to a human triage queue.
  • If the Score is angry, alert the duty manager on Slack.

Jev handles the fuzzy perception; plain code handles the business logic.

jev-use-case

Trying the API

The docs show a single lightweight endpoint. Here's a raw cURL request:

curl https://api.typesafe.ai/v1/systemone \ -H "Authorization: Bearer $TYPESAFE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "state": "Help! My payouts have been failing for 3 days.", "model": "jev-latest", "questions": { "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"} } }'

The response returns clean, typed data:

{ "is_urgent": { "type": "noul", "noul": 0.92 } }

Python SDK Implementation

You can install the official client (pip install typesafe-sdk) and run multi-question classification in just a few lines:

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient # 1. Provide the unstructured context (the State) state = { "message": "I was double charged on my invoice and our payroll run is blocked!", "account_tier": "enterprise", } # 2. Define your schema upfront questions = { "intent": Choice( instructions="What is the customer's primary issue?", criteria={ "billing": "Invoices, charges, refunds, or payment issues.", "technical_bug": "Software defects or system crashes.", "feature_request": "Asking for new functionality.", "other": "General inquiries or uncategorized messages.", }, ), "is_urgent": Noul( instructions="Does this issue block core business operations or require immediate action today?" ), "sentiment": Score( instructions="Rate the user's emotional tone.", scale=["calm", "frustrated", "hostile"], ), } # 3. Parallel evaluation in a single round-trip with TypeSafeClient() as client: result = client.system_one(state=state, questions=questions) intent = result.choices["intent"] print(f"Category: {intent.choice} (Confidence: {intent.confidence:.2f})") urgency = result.nouls["is_urgent"] print(f"Urgent probability: {urgency.noul:.2f}") sentiment = result.scores["sentiment"] print(f"Tone: {sentiment.score}")

Tip: Always include an "other" option in your Choices. If you don't, Jev is forced to pick the closest category even when the input is irrelevant.

Vercel AI Gateway also offered free tier access for Jev during launch week, with Netlify AI Gateway supporting it as well.

The Limits

Before replacing every router in your stack, keep its boundaries in mind:

  • No explanations: you get answers and numbers, not reasoning steps or chain-of-thought audits.
  • Text only (for now): it doesn't accept images, video, or audio natively. You must OCR or transcribe your inputs first.
  • No external browsing: it cannot search the web or fetch external URLs; it only knows what you pass in the state.
  • Valid output ≠ correct truth: while the schema cannot break, the classification can still be incorrect if the instructions are ambiguous.
  • Token limits: the state plus your questions must comfortably fit within early-access context limits (around 32k tokens).
  • Closed API: while open-source experiments like openjev, SemIf, and kev have surfaced, the core Jev model is hosted and proprietary.

One Brain, Many Reflexes

What makes Jev exciting isn't that it's smarter than GPT-4 or Claude—it's that it isn't trying to be.

The design pattern that makes the most sense moving forward is one brain, many reflexes:

  • Use a large, expensive LLM when you need creative writing, deep multi-step synthesis, or nuanced code generation.
  • Use fast, cheap System 1 classifiers for everything else: routing, filtering, content safety, tool gating, and intent extraction.

Once models like Jev get native vision ("giving them eyes"), browser and computer-use automation will become significantly faster and cheaper. It won't take long before major model providers launch their own dedicated System 1 endpoints.

Eventually, models like Jev will simply vanish into our infrastructure stack—because when a decision is this fast and this cheap, using an LLM to answer "yes or no" makes no sense at all.

3 Claps

Share this article

Help others discover this content

Thanks for reading! 👋

More Articles
3