Jev AIDecision ModelsAI Agents

Jev: The AI Model That Costs Almost Nothing and Makes Decisions

Jev makes typed AI decisions for almost nothing. Learn how it differs from LLMs, where it works, and which industries may adopt it first.

Yaroslav Dobroskok9 min read
Get new articles by email

One dollar buys roughly 23,800 Jev decisions, assuming 1,000 input tokens per call. At that price, the interesting question becomes where you would put them.

Between every tool call in a coding agent? In every incoming webhook? Before each retry in a failing CI pipeline? Suddenly, decisions you would have left to brittle keyword rules look worth revisiting.

There is a catch that makes this model particularly interesting to developers: it cannot write you an answer. You have to define the answers it is allowed to give.

Jev AI, a new model from TypeSafe AI, is built specifically for that job. It reads a state, answers predefined questions and returns typed choices, scores or probabilities that software can use directly. It does not chat, write or explain.

Browser Use's recorded flight-search demo, shown at original speed. Jev selects actions and targets; a separate model supplies text when needed. Browser Use — source and measurements

Jev pricing: $0.042 per Million Input Tokens

The published price behind that calculation is $0.042 per million input tokens, with output free. Gateway fees and the size of each request still matter. But the starting cost is low enough to reconsider where semantic decisions belong in software.

Jev vs LLM: How a Decision Model Differs

TypeSafe calls Jev its first System One model, borrowing the name from the fast, intuitive mode of thought in Daniel Kahneman’s Thinking, Fast and Slow.

Jev behaves like a classifier with the semantic understanding we associate with modern language models.

You give Jev two things:

  • State: the information it should inspect, such as a customer message, agent trace or list of available actions.
  • Questions: predefined decisions about that state.

Jev supports three question shapes. A Choice picks from options you define. A Score places the state on a scale. A Noul returns a number from 0 to 1 for a yes-or-no proposition. The response includes probabilities and confidence, so your application can act above a threshold and fall back when the answer is uncertain.

An LLM works differently. It generates a string one token at a time. This is also where Jev departs from structured outputs, the feature most developers reach for today. JSON mode and schema-constrained decoding still ask a generative model to produce tokens, then constrain and validate the result: you get a typed envelope around a generative process. Structured outputs shape the answer; Jev never generates one. That flexibility is exactly what makes LLMs useful for writing, coding and open-ended reasoning—but it is unnecessary when your program needs one of five known actions.

Jev gives up text generation. In return, it produces only values that fit the schema you defined.

That does not mean it is always correct. TypeSafe’s “zero hallucinations” claim is best understood narrowly: Jev cannot invent an option outside your schema or return malformed prose instead of a typed answer. It can still choose the wrong valid option. Confidence thresholds, evaluations and fallback paths remain part of the production design.

Jev is therefore not a replacement for Claude or GPT. A useful split is:

Jev decides. An LLM creates. Code enforces the policy.

What can you build with it?

The first experiments show why that split matters. For a longer, illustrated list, see 12 Jev use cases for coding agents and beyond.

Route support tickets in one call

Imagine receiving this message: “I’ve been trying to connect Stripe for three days. I’m losing sales. Please help ASAP.”

One Jev request can decide the department, score the customer’s frustration and estimate whether the issue is urgent. TypeSafe’s published example returned “billing” with 0.84 probability, an urgency value of 0.999 and a frustration score—all from the same state, in parallel.

Your application can route high-confidence cases immediately and send ambiguous ones to an LLM or human operator.

Put a decision at every step of an AI agent

Agents constantly make small fuzzy decisions: which tool to call, whether an action is risky, whether the result meets the acceptance criteria and whether the next request needs a stronger model.

The Nexus Agent team tested Jev as a guardrail on 469 held-out cases. Jev made a decision on 58% of them and fell back to existing rules on the rest. Where it did decide, it made 5 mistakes in 274 decisions, compared with 22 in 316 for Claude Sonnet 5 and 47 in 342 for Claude Haiku 4.5. The resulting system had a similar total number of errors to the LLM alternatives. Nexus reported about 9 times lower median latency and 79 to 214 times lower cost per answer, though its LLM timings included CLI startup.

Those are one team’s results on one workflow, not a universal benchmark. They are still a useful illustration of the right architecture: let the fast model handle cases above a confidence threshold, and preserve the old path for everything else.

Make browser agents feel interactive

Browser Use published a hybrid browser agent in which Jev chooses the next page action and a small generative model writes text only when needed. Its measured Google Flights run finished in 7.073 seconds. The repository describes the limits clearly: this was one task on one browser profile, not a general reliability benchmark.

Still, the demo makes the latency argument visible. An LLM that pauses for several seconds cannot sit comfortably inside an interactive control loop.

Why is Jev so fast and cheap?

The biggest reason is architectural restraint.

An LLM must preserve the ability to produce almost any string. It generates that string sequentially, token after token. Output can be much longer than the decision the software actually needs, and output tokens usually cost more than input tokens.

Jev does not generate a string. It evaluates predefined questions and produces their answers in parallel. TypeSafe says its stack combines a new architecture, a parallel sampler and a training method called Reinforcement Learning for Calibrated Decisions (RLCD). The training objective rewards both the correct decision and an honest estimate of uncertainty.

This is also why asking several narrow questions in one request is attractive. TypeSafe reported that answering 13 questions together was 10 times faster and 12.2 times cheaper than making 13 separate calls.

The vendor reports 70–500 milliseconds end to end. Independent measurements can be slower: Nexus measured a 680-millisecond median across 36,218 calls from Europe, with large contexts and fresh connections. With connection reuse, its single calls fell to 230–350 milliseconds.

TypeSafe’s largest benchmark claims—193.6 times faster and 444.6 times cheaper than LLM workflows—should be treated as vendor results. TypeSafe itself says these numbers are likely at the high end of real-world gains, and that its workflow authors may have introduced bias. The underlying price and absence of generated output are easier to verify than broad claims about intelligence.

Where will Jev be adopted first?

The first adoption will likely happen where the decision is narrow, repeated often and easy to evaluate.

Software engineering is where I would look first. Software already has typed interfaces, enumerated actions and explicit states. A build failure might lead to RETRY, INVESTIGATE or ESCALATE; an issue might belong to one of six services; a coding agent might choose between searching, editing and running tests. The answer space is already designed. The missing piece is interpreting messy logs, descriptions or execution traces well enough to choose within it.

That makes Jev a promising fit for the semantic decisions between deterministic components. For example, an illustrative CI workflow could classify a failure as a likely infrastructure problem, a test regression or an unknown, then let ordinary code enforce retry limits and route the result. A state machine can expose only transitions valid from its current state; Jev supplies the judgment about which of those transitions fits the evidence.

The interface can be deterministic while the judgment remains probabilistic. Jev does not replace a compiler, a permission check or an exact business rule. Its appeal is that it can fit the interfaces developers already use, with confidence thresholds and a fallback for ambiguous inputs.

AI agent infrastructure is the clearest fit. Tool selection, model routing, completion checks and action-risk gates happen on nearly every agent turn. Saving seconds and cents at each step compounds quickly.

Customer support, recruiting and operations should follow. Recruitly says it put Jev into production on launch day for reply classification, support urgency, job-category selection, candidate scoring and knowledge-base checks. Its low-confidence cases still go through the previous LLM path, which limits migration risk. This is a self-reported case study, but the implementation pattern is sound.

High-volume content and data pipelines are another natural entry point: document classification, moderation, deduplication, product taxonomy and RAG relevance checks. These tasks need semantic understanding, yet usually produce a label rather than prose.

Interactive agents and real-time interfaces may be the most transformative category. Browser automation, games and responsive assistants have been held back by multi-second model calls. Sub-second decisions change the user experience.

Regulated and safety-critical systems will move more slowly. A probability is useful, but it is not an audit trail or a correctness guarantee. Healthcare, finance, industrial control and autonomous vehicles require domain evaluations, deterministic safeguards and human escalation—not a launch-week demo.

The bigger idea

Jev is interesting because it challenges a lazy assumption in AI product development: that every intelligent operation should be a prompt to a general-purpose LLM.

Many software problems are not generation problems. They are thousands of tiny decisions hidden inside workflows.

If Jev’s quality holds up outside early examples, the winning architecture will rarely be “replace the LLM.” It will be a hybrid system: code for hard constraints, Jev for fast semantic decisions and an LLM for open-ended reasoning and creation.

The model is only days old, access is still early, and most public evidence comes from TypeSafe or enthusiastic first adopters. We should be skeptical of sweeping benchmark claims.

But the economic shift is already clear. At $0.042 per million input tokens, a semantic decision can become cheap enough to place almost anywhere in software. When intelligence stops being a special event and becomes another function call, developers start asking a much more interesting question:

What would we let software decide if the decision cost almost nothing?

Running a software company or leading an engineering team?

Take the free automated AI Adoption Healthcheck: 3 minutes, and you get your company's AI adoption score with next steps.

Get the next article in your inbox

Practical notes on AI tools and engineering workflows. I'll email you when a new article is published.

Only new article updates. Unsubscribe anytime.

How we use your email