For the last few years, most AI systems have been built around one basic idea:
give the model some input and ask it to generate an answer.
That works well when a person needs text, code, an explanation or a conversation.
But software often needs something much simpler.
Which team should receive this request?
Which tool should the agent call?
Should this action continue or stop?
Is this output acceptable?
Which model should handle the next step?
Those are not writing problems.
They are decision problems.
A new group of AI systems is now being designed specifically for that kind of work.
TypeSafe’s Jev, OpenAI’s Decisions API, and AWS’s open-source Strands Decider all point in the same direction:
sometimes software does not need an AI that talks. It needs an AI that chooses.
Why use a full language model for every decision?
Consider an AI support agent.
It receives this request:
“My payout has failed for three days.”
The system may need to decide whether the request belongs to:
- billing
- sales
- retail support
A large language model can make that decision.
You can prompt it:
Return only one of these three values.
You can also ask it to produce structured JSON.
That works.
But the model is still fundamentally a text generator.
It processes the request, predicts tokens and generates an answer even though the software already knows every legal output.
The application does not need prose.
It needs:
billing
plus, ideally:
how confident are you?
That is the gap decision-oriented systems are trying to fill.
Accessible text alternative for this figure
Two flows side by side. Generative AI: input goes to a model, which produces open-ended generated output, for example an explanation of why a payment failed. The output box is drawn with a dashed outline because its shape is not fixed in advance. Decision AI: state and allowed choices go to a decision system, which returns a selected option and a confidence or probability where supported, for example billing, sales or support. Every step has a text label, so no meaning depends on colour.
Jev made the distinction explicit
TypeSafe introduced Jev in September 2026 as its first System One Model.
We covered the model itself in an earlier TechiesJournal article, so this piece focuses on the wider category.
The company describes the interface simply:
unstructured state in → typed probabilistic decisions out
Instead of generating text, Jev receives some state and a set of developer-defined questions or options.
It returns structured choices and probabilities.
TypeSafe says Jev uses a new architecture, a parallel sampler and a training approach called Reinforcement Learning for Calibrated Decisions, or RLCD. The goal is not merely to choose the right option, but to make the reported probabilities useful as estimates of uncertainty.
That means an application could define rules such as:
confidence above 0.95 → continue automatically
confidence between 0.70 and 0.95 → perform another check
confidence below 0.70 → ask a person
This is very different from treating a model response as trustworthy simply because it sounds confident.
But one limitation must be clear:
calibrated does not mean always correct.
A model can still choose the wrong answer.
Calibration means that across many similar decisions, its probabilities should correspond reasonably well to how often those decisions are correct.
Independent research published shortly after Jev’s launch found encouraging results across a large collection of classification and reasoning datasets, including useful probability calibration, while also finding weaker performance in areas such as noisy labels, fine-grained classification and some rubric-based judgments.
So Jev’s early evidence is promising, but decision models do not eliminate evaluation.
OpenAI reached the same problem from a different direction
At DevDay 2026, OpenAI announced the Decisions API.
The interface is intended for tasks such as:
- classification
- request routing
- choosing an agent’s next action
- repetitive business decisions
The similarity to Jev is obvious.
The developer defines a decision space rather than asking for an open-ended response.
But there is an important technical difference.
OpenAI says the Decisions API is powered by GPT-6 Luna.
OpenAI has not publicly described it as a separate Jev-style model architecture or disclosed a Jev-like training method dedicated to calibrated decisions.
So today it is safer to describe the Decisions API as:
a decision-oriented service built on OpenAI’s broader model platform
rather than:
a specialized decision model equivalent to Jev.
That difference could matter.
Using Luna potentially gives OpenAI access to broader language and multimodal capabilities.
A more specialized system such as Jev may have advantages in efficiency, calibration or predictable machine-oriented outputs.
We do not yet have enough public evidence to determine where that trade-off lands.
The Decisions API is also still in limited preview, so detailed independent testing remains sparse.
AWS took the specialization approach further
AWS’s Strands team then released Strands Decider, an open-source decision model intended for agentic workflows.
Its architecture is unusually transparent.
The project starts with a pretrained Qwen3.5-2B language-model body, removes the language-generation head, and replaces it with a small decision head.
That means the model no longer generates arbitrary text.
It scores predefined choices directly.
The first released model has about 1.9 billion parameters, and AWS reports a median response time of roughly 115 milliseconds on an RTX 3090 for its reference setup. It can also run locally on Apple Silicon and CPU-based systems.
Its intended uses include:
- model routing
- tool selection
- tool-argument checking
- triage
- guardrails
- evaluations
- hybrid agent workflows
The last one is particularly interesting.
AWS suggests using a larger language model for genuinely difficult reasoning while letting a smaller decision model handle routine choices.
That could become an important architecture pattern.
The future agent may use more than one kind of intelligence
Many agent systems today look roughly like this:
LLM → choose tool → LLM → inspect result → LLM → decide next step → LLM → evaluate output
Every intelligent step goes through the same general-purpose model.
Decision models suggest another design:
decision model → routine routing
decision model → tool selection
large model → difficult reasoning
decision model → output check
human → high-risk or uncertain case
Accessible text alternative for this figure
Top: the traditional pattern, in which one general-purpose large language model makes every step: choose tool, inspect result, decide next step, evaluate. Bottom: a possible layered pattern with five roles. A decision model handles routine routing. A decision model handles tool selection. A large language or reasoning model handles difficult reasoning. A decision model handles routine evaluation. A human handles uncertain or high-consequence decisions. The figure is labelled as an architectural possibility, not an established universal best practice. Each role has a text label, so no meaning depends on colour.
That separation could reduce:
- latency
- inference cost
- unnecessary generation
- variability
It could also make some decisions easier to audit because the permitted outputs are defined before the model runs.
But it introduces another responsibility:
developers need to decide which decisions are safe to constrain.
Some decisions should not be compressed into a menu
Decision models work best when the possible outcomes are known.
For example:
Which support queue?
Which tool?
Is this transaction suspicious?
Does this answer satisfy the policy?
But some tasks are genuinely open-ended.
A model may need to:
- investigate
- explain
- write
- synthesize evidence
- propose a plan
- discover an option the developer did not anticipate
If the correct answer is not in the predefined choice set, a decision model cannot invent it.
That is both a strength and a limitation.
The output cannot wander outside the permitted structure.
But the system designer is responsible for making sure the structure itself is adequate.
“Cannot hallucinate” needs careful wording
TypeSafe says Jev cannot hallucinate because it does not generate arbitrary strings.
There is a useful idea behind that statement, but it can easily be misunderstood.
Suppose the allowed outputs are:
- approve
- reject
A decision model cannot suddenly return:
“Send this to another department.”
That is true.
But it can still incorrectly return:
approve
when the correct decision was:
reject.
So a more precise description is:
A constrained decision model cannot generate an out-of-schema answer, but it can still make an incorrect in-schema decision.
That is why calibration and evaluation matter so much.
Why developers are paying attention
AWS distinguished engineer Marc Brooker told TechCrunch that the motivation came from agent workflows where customers did not need the full cost or capability of an LLM for every step.
He described decision models as particularly useful for questions such as:
what should the workflow do next?
and highlighted the combination of constrained answers, confidence scores, lower latency and potentially lower cost.
Independent developer interest around Jev has followed a similar pattern: routing, classification, robotics and agent control are among the recurring use cases.
This suggests the real demand is not for another chatbot.
It is for small pieces of intelligence embedded inside ordinary software control flow.
Are these just classifiers with new branding?
That is a reasonable question.
Traditional machine-learning classifiers have handled structured decisions for years.
A fraud classifier, sentiment model or routing model can already return a probability.
The difference these newer systems are trying to offer is generality.
A traditional classifier normally needs:
- a defined task
- labeled data
- training
- deployment
- maintenance
A decision model aims to accept a new natural-language decision at runtime without training a separate classifier for every individual task.
In that sense, it sits somewhere between:
traditional classifier
and
general-purpose LLM.
AWS explicitly describes Strands Decider this way: useful for problems where a traditional classifier would require too much task-specific training, but a full LLM is unnecessary.
That may turn out to be the category’s most useful definition.
The biggest opportunity may be inside agents
This development connects directly to the move toward longer-running AI agents.
An agent may make hundreds of small decisions while completing one task.
If every choice requires a frontier model call, the costs and delays accumulate.
Consider:
- Which document should I retrieve?
- Is the search result relevant?
- Which API should I call?
- Are these arguments valid?
- Did the tool produce a usable result?
- Should I retry?
- Should I escalate?
- Is the final answer sufficiently grounded?
Not all of those require frontier-level reasoning.
If specialized decision models can handle the routine choices reliably, larger models can be reserved for the parts where deeper reasoning actually adds value.
That may make agent systems cheaper and faster without necessarily making them less capable.
My Perspective: intelligence does not need one interface
For years, generative AI has encouraged a simple assumption:
if software needs intelligence, call an LLM.
Jev, OpenAI Decisions API and Strands Decider challenge that assumption.
They suggest a more layered future.
Use a language model when software needs to generate, reason or explore.
Use a decision system when software already knows the possible outcomes and needs help choosing between them.
Use deterministic code when the answer should not require AI at all.
And use a human when uncertainty or consequence makes automation inappropriate.
The important question is therefore not:
Will decision models replace LLMs?
Probably not.
It is:
How much of today’s LLM workload never needed language generation in the first place?
If the answer is substantial, decision-oriented AI could become an important layer inside future software and agent systems.
Not because machines have stopped needing intelligence.
Because intelligence does not always need to talk.
Sources and Further Reading
Source review: 2 October 2026.
- TypeSafe AI — Introducing System One Models & Jev. Primary launch material covering Jev’s architecture, decision interface and RLCD training claims.
- Strands Labs — Strands Decider. Open-source model architecture, evaluation, local deployment information and agent-workflow use cases.
- Evaluating and Benchmarking the System One Model Jev. Independent evaluation across 37 datasets, including accuracy and calibration findings.
- Calibrated Decisions at Scale. Applied research using Jev for large-scale structured coding and human-review gating.
- OpenAI — DevDay 2026 Announcements and Developer Resources. Official summary of the Decisions API announcement and its limited-preview status.
- TechCrunch — Amazon releases its own Jev clone as decision models flood the web. Industry commentary and AWS engineering perspective on decision models in agentic workflows.
