ProductOS
Jev and the unbundling of AI: why agents need a decision layer

Jev and the unbundling of AI: why agents need a decision layer

Heemang Parmar

Heemang Parmar

Published ·14 min read

Most AI agents are using a frontier language model for decisions that should cost a fraction of a cent.

Which tool should run next? Which model should handle this task? Is this support request urgent? Does the evidence support the claim? Has the agent actually completed the work?

We usually send all of these questions to the same general-purpose LLM that writes the final response. It works, but the architecture is becoming difficult to justify. We are paying for open-ended generation when the application only needs a bounded decision.

Jev is an early attempt to separate those two jobs.

TypeSafe AI introduced Jev on 15 September 2026 as its first “System One” model. It does not generate paragraphs, code, or explanations. It receives a state and a set of typed questions, then returns constrained answers with probabilities.[1][2]

It is only 12 days old as I write this, so I would not treat it as settled infrastructure. But the idea behind it matters beyond one model or company.

AI is starting to unbundle.

The best model for writing may not be the best model for routing. The best model for reasoning may not be the best model for checking a policy. The best model for a conversation may be excessive for deciding whether a ticket belongs to billing or technical support.

That shift could change how we design agents.

What Jev is

Jev is a decision model.

A normal LLM accepts context and generates a sequence of tokens. Even when we ask it for JSON, it is still generating text. The application then parses that text, validates the schema, checks whether the values are permitted, and decides what to do next.

Jev removes the text-generation step from the interface.

A request contains:

  • A state, supplied as text or structured JSON
  • One or more named questions
  • The permitted shape of each answer

The response contains typed values and probability distributions that application code can inspect directly.[3]

The current API exposes three question types.

Choice

Choice selects one option from a closed set.

A support system could ask:

Which team should handle this message?

billing: charges, invoices, or refunds
technical: bugs, outages, or integrations
account: login, plan, or profile changes

Jev returns the selected option, a probability for every option, and a confidence score. Choice supports up to 255 options.[3]

The application does not need to extract billing from a paragraph or handle an invented category such as customer success. The answer must fit the declared set.

Score

Score places an input on an ordered rubric.

For example:

How severe is this issue?

0: cosmetic problem
1: workflow degraded
2: core workflow blocked
3: security, data, or production impact

The result can fall between levels because it is calculated from the probability distribution. This works for urgency, risk, quality, frustration, lead fit, or any judgment where we can write clear levels.

Noul

Noul is a yes-or-no proposition expressed as a probability from 0 to 1.

Examples:

  • Does this message request a refund?
  • Does this source directly support the claim?
  • Does the proposed tool call have irreversible side effects?
  • Does this task require a human reviewer?

A result near 1 is a strong yes. A result near 0 is a strong no. A result around 0.5 signals uncertainty.

The name may be unfamiliar, but the programming model is simple. Noul maps to an if. Choice maps to a branch. Score maps to a threshold or ranking.

Why this interface matters

A four-stage AI decision architecture separating generative models, Jev, deterministic code, and human review
One possible production architecture: generative models create, Jev evaluates bounded decisions, code enforces policy, and humans review uncertain or high-risk cases.

Typed output is useful, but it is not the most interesting part.

The larger change is architectural: code owns the workflow again.

Many agent systems put the workflow inside a prompt. The model decides what the task means, what to do next, which tool to call, whether the result is good enough, and when to stop. This is flexible, but it gives one probabilistic component a large amount of responsibility.

A decision layer creates a narrower boundary:

Generative model creates or reasons
                ↓
Decision model evaluates a bounded question
                ↓
Deterministic code applies policy
                ↓
Proceed, escalate, ask a person, or use a stronger model

LangChain describes a similar pattern in its work with Jev and LangGraph: code retains the workflow, decision models handle narrow judgments, and uncertain cases move to an LLM or human reviewer.[7]

This separation gives us a few practical benefits.

First, the answer space is explicit. A model can choose the wrong option, but it cannot silently create a new one.

Second, uncertainty becomes part of the interface. Code can treat a 0.58 result differently from a 0.97 result.

Third, the same state can be evaluated against several independent questions in one request. A support message can be classified for queue, urgency, frustration, policy risk, and human escalation together.[3]

Fourth, the expensive model only runs where it adds value. A routine high-confidence route can remain cheap. An ambiguous or high-stakes case can move to a stronger model or a person.

This is closer to normal software engineering than a fully open agent loop.

The current economics

TypeSafe currently lists jev-1.13.0 as the official model. The jev-latest alias points to it, although that alias will move when a new stable model ships.[4]

The published price is $0.042 per million input tokens, with free output. At that rate, 1,000 requests containing 1,000 input tokens each cost about $0.042. One million such requests cost about $42.

The official limits at the time of writing are:

  • 64k total tokens per request
  • 32k tokens for the state plus the longest question
  • Text-only input
  • 250,000 tokens per second
  • 1,200 requests per minute

TypeSafe notes that launch-period rate limits may change.[4]

The company reports 70 to 500 millisecond response times. Its launch evaluation also reports gains up to 193.6 times faster and 444.6 times cheaper than LLM-based workflows.[1]

Those numbers need context.

TypeSafe says the reported gains are likely toward the high end of real-world improvements. It also notes that members of its model-capabilities team created the workflows and that the reference probabilities came from frontier models rather than human ground truth.[1]

I read those as promising engineering results, not a universal benchmark.

What independent testing shows

The most useful independent evaluation I found tested Jev 1.13.0 on 400 requests containing 2,000 individual decisions.

DecisionEval reported:[6]

  • 74.0% overall accuracy
  • Expected calibration error of 0.045
  • P50 latency of 687 milliseconds
  • P95 latency of 777 milliseconds
  • 86.4% accuracy when confidence was at least 0.7, covering 57.6% of cases
  • 93.7% accuracy when confidence was at least 0.9, covering 24.5% of cases

The headline accuracy is not the main lesson.

The coverage curve is.

At a high confidence threshold, Jev handled fewer cases but was more accurate on those it retained. That is the production pattern I would expect:

  • Automate clear and reversible cases
  • Escalate ambiguous cases
  • Require human review for high-impact actions

A separate test by TrueStandard evaluated 108 claim-and-source pairs across six domains. Jev was slightly more accurate and materially cheaper than Gemini Flash Lite and Claude Haiku in that test, especially on adversarial near-miss claims. But its calibration error was similar to the two chat models, not clearly better.[10]

The sample was small, and the author is explicit about how the conclusion changed as more examples were added. That is useful in itself. Small evaluations can create confident stories that disappear with a larger dataset.

The practical conclusion is narrower:

Jev looks competitive for constrained judgments. It is very cheap. It provides a cleaner machine interface. Its confidence may be useful for routing. None of this removes the need to test it against our own traffic.

“Cannot hallucinate” needs a correction

TypeSafe says Jev cannot hallucinate because it cannot return a value outside the supplied schema.[1]

That is true in a specific engineering sense.

If the allowed options are billing, technical, and account, Jev cannot return legal. It cannot emit malformed JSON or replace the answer with a paragraph. This removes an entire class of integration failures.

But a schema-valid answer can still be wrong.

Jev can confidently select billing when the correct route is technical. It can assign high urgency to an ordinary request. It can decide that evidence supports a claim when the support is weak.

TrueStandard puts the distinction clearly: Jev cannot invent an option outside the declared set, but it can choose the wrong permitted option.[10]

For production teams, this means:

  • We may no longer need to validate the shape of the output
  • We still need to evaluate the quality of the decision
  • Confidence is a routing signal, not authorization
  • Local outcome data matters more than a vendor’s calibration claim

Calibration is distribution-specific. A model can be well calibrated on the data used during development and poorly calibrated on a team’s customer messages, taxonomies, tool traces, or policies. Prefactor’s analysis makes the same point: thresholds become trustworthy only after reported confidence is compared with real outcomes on representative production traffic.[11]

Where Jev fits inside an agent

I see five strong control points.

1. Skill and tool selection

An agent often receives a large list of tools or skills and asks a frontier model to choose one. As the catalogue grows, the prompt becomes larger and tool selection becomes less reliable.

A decision model can evaluate the task against a bounded shortlist:

  • Choice: which skill best matches?
  • Noul: does any candidate actually apply?
  • Confidence gate: should the selection be accepted or passed to the main model?

TypeSafe has published a cookbook that applies this pattern to a catalogue of 182 agent skills.[4]

This is one of the clearest early experiments for any agent with a growing tool or skill catalogue. It can reduce missed specialist workflows without changing the main conversational model.

2. Model routing

Not every request needs the same model.

A decision layer can choose among routes such as:

  • Fast model
  • Strong reasoning model
  • Coding model
  • Vision model
  • Multiple subagents
  • Human clarification

The output should not directly call the chosen provider. Code should check that the model is available, enforce cost and policy limits, and execute the route.

The benefit is not only lower cost. It is making model-selection behavior visible and measurable.

3. Semantic tool guards

Deterministic security rules are necessary, but they cannot understand every proposed action.

A semantic guard can inspect:

  • Tool name
  • Sanitized argument summary
  • Expected side effects
  • Reversibility
  • Applicable policy
  • Existing safeguards

It can then return allow, ask_user, review, or deny.

This must remain advisory. A decision model should never override permissions, user approvals, or deterministic deny rules. It adds judgment around the policy boundary. It does not become the boundary.

4. Completion verification

Agents often stop because the generated response feels complete, not because every acceptance criterion has evidence.

Before an agent reports completion, a decision layer can evaluate:

  • Objective
  • Acceptance criteria
  • Test results
  • Tool evidence
  • Deployment status
  • Read-back verification
  • Known failures

The output might be complete, verification_missing, incomplete, or human_review.

This is particularly useful when an agent has already spent most of its context on implementation. A compact independent check can catch missing verification before a confident final message reaches the user.

5. Research and citation checks

Given a claim and a source passage, Jev can judge whether the passage supports the claim.

This can act as a high-volume pre-publication check for research, reports, proposals, and generated content. It does not establish whether the source is itself correct or current. It only evaluates the relationship between supplied evidence and the claim.

That is still useful. Many citation failures are not invented URLs. They are real sources attached to claims the source does not actually support.

Browser agents are an early example

Browser Use has published an experimental jev-ultrafast agent.

The browser is converted into a dynamic list of valid operations and interactive elements. Jev chooses an operation and target. A small LLM generates text only when the operation requires typing.[8]

In a small Google Flights test, the project reported a reduction in median task time from 9.450 seconds to 7.092 seconds and a reduction in browser protocol calls from 1,092 to 101. The test covered three repeats per version on one task, and the authors explicitly say it is not a general reliability benchmark.[8]

The result is still interesting because it demonstrates the architecture clearly.

A browser action feels open-ended, but at any given moment the page offers a finite set of valid controls. Once the state is converted into that bounded action space, tool selection becomes a decision problem.

Where Jev should not be used

TypeSafe’s own limitations page is unusually direct, which I appreciate.

Jev 1.13 can struggle with:[5]

  • Literal phrasing and implied conditions
  • Counting and numeric precision
  • Date and time comparison
  • Multi-step indirection
  • Large states containing irrelevant detail
  • Adversarial content that tries to influence classification
  • Contradictory instructions and criteria
  • Structural assumptions across separately asked questions
  • Text generation

The safe design response is simple.

Use code for arithmetic, dates, permissions, exact matching, inventory, pricing, and policy enforcement.

Use retrieval or filtering before sending context.

Ask one focused judgment per question.

Use a generative model for writing and open-ended reasoning.

Keep a human path for uncertain or consequential cases.

Jev’s primary training language is English. Other languages are supported but currently have lower accuracy, so multilingual production traffic needs separate evaluation.[4]

The model is also proprietary and hosted. TypeSafe says customer requests and responses are not used for training, and enterprise customers can request zero data retention. Vercel AI Gateway also documents per-request Zero Data Retention and No Training options for Jev.[4][9]

Data sensitivity and provider terms still need to be reviewed before sending production content.

How I would test it

I would not begin with a platform-wide integration.

I would run a shadow pilot on one bounded workflow where the team already knows the correct historical outcomes. Skill selection, support triage, catalogue matching, and completion verification are all reasonable candidates, depending on where clean labels already exist.

The pilot would follow this sequence:

  1. Collect 300 to 500 representative historical examples, including difficult and ambiguous cases.
  2. Write narrow questions with explicit options and boundary cases.
  3. Pin the versioned model ID instead of using the moving jev-latest alias.[4]
  4. Run Jev in shadow mode without allowing it to take action.
  5. Store the complete probability distribution, confidence, model version, question version, latency, cost, and correct outcome.
  6. Plot accuracy against coverage at different thresholds.
  7. Review every high-confidence error.
  8. Set separate thresholds for each question and primitive.
  9. Auto-route only low-risk, reversible cases after the held-out results clear the required error bar.
  10. Re-test whenever the model, question wording, or input distribution changes.

The benchmark that matters is not whether Jev beats a frontier model on a public chart.

It is whether a specific decision can be automated on our data at an acceptable error rate, latency, and cost.

The larger shift

For the past few years, the default AI architecture has been one powerful model connected to many tools.

That was the fastest way to discover what agents could do. It is probably not the final production architecture.

We are beginning to separate intelligence into components:

  • Generative models for creation and explanation
  • Reasoning models for difficult planning
  • Vision models for visual understanding
  • Embedding and reranking models for retrieval
  • Decision models for bounded semantic judgments
  • Deterministic code for policy, computation, and side effects

The work then moves from finding one model that does everything to designing the system that gives each model the right job.

That is the part I find most relevant.

Jev may become an important model, or another provider may build a better decision layer. Either way, the interface makes sense. Agents need a way to make cheap, typed, observable judgments without turning every branch into another open-ended generation call.

The right starting point is not broad adoption. It is one measurable decision, running quietly in shadow, until the probabilities earn the right to influence the workflow.

If you are building production agents, I would be interested to compare where your stack still uses a frontier model for a bounded decision.

Sources

[1] https://typesafe.ai/blog/introducing-system-one-models-and-jev [2] https://docs.typesafe.ai/introduction [3] https://docs.typesafe.ai/api [4] https://docs.typesafe.ai/models [5] https://docs.typesafe.ai/model-jaggedness/jev-1.13 [6] https://decisioneval.dev/models/typesafe-jev [7] https://www.langchain.com/blog/building-prod-with-jev-and-langgraph [8] https://github.com/browser-use/jev-ultrafast [9] https://vercel.com/changelog/typesafe-ai-jev-now-available-on-ai-gateway [10] https://truestandard.ai/blog/jev-accuracy-tested [11] https://prefactor.tech/blog/jev-calibrated-confidence-is-not-correctness

Heemang Parmar

Heemang Parmar

CS engineer and IIM Lucknow MBA. Built products across enterprise and AI for 10+ years. Founded ProductOS to give every PM and founder the leverage of a full product team. Writes about AI product development, PRDs, and building with agents.