CODEHANCEBlog
Browse allAbout
← All posts
Building with AI·7 October 2026·12 min read

Jev, OpenAI, and the Shift Towards Decision AI

I started by asking what Jev actually built. A four-subagent RAG example brought the opportunity into focus: useful decisions need useful context and software that knows what to do next.

Kingsley Ijomah

Kingsley Ijomah

AI Adoption Lead

Updated 7 October 2026

Start the discussion
Share

Share this article

LinkedInFacebookXWhatsAppEmail
Evidence cards feed a small decision component that branches into retrieval, an answering model and human review.

On this page

10 sections, in order. Jump straight to the one you need.

  1. 01What Did TypeSafe Actually Build with Jev?
  2. 02How Is This Different from AI We Already Have?
  3. 03Where Does OpenAI’s Decisions API Fit?
  4. 04Context Comes First
  5. 05A Decision Model Needs a System Around It
  6. 06Where Jev Could Fit Inside an AI Agent
  7. 07What Could This Change for People and Teams?
  8. 08How to Try Decision AI Today
  9. 09What Would Make This Shift Matter?
  10. 10Sources

Follow the work as it becomes practical.

Get new field notes, hands-on examples, and build videos in your inbox.

Practical AI notes, no noise. Unsubscribe whenever you like.

I started looking at Jev after coming across OpenAI’s Decisions API video. Before writing about a new trend, I wanted to understand what had actually changed. Had TypeSafe trained a different model? Was OpenAI offering the same thing? Could an ordinary language model already do this?

The question that made it concrete for me was about RAG: if I decompose a question into four subquestions and send four subagents to retrieve evidence, where would Jev fit?

Following that example changed the focus of the article. The interesting possibility is a model that can judge evidence at specific points in a workflow, quickly enough and cheaply enough to make those checks practical. But the evidence still has to reach it.

Decision AI depends on the context we provide and the software we build around it. I’m exploring that design here through proposed experiments, rather than reporting a benchmark I have run. That distinction matters when a promising interface is being presented as a change in how we build.

What Did TypeSafe Actually Build with Jev?

TypeSafe AI is the company behind Jev. Its founders are CEO Diogo Almeida, COO Sasha Sheng and CTO Erik Gafni. Their backgrounds span OpenAI, Google Brain, Meta/FAIR and production AI systems, according to the company’s team page.

Almeida’s research connection is particularly relevant. He is a co-author of the 2022 InstructGPT paper, which describes training language models to follow instructions using human feedback. Someone who contributed to that work is now pursuing a different interface for software.

TypeSafe released Jev in early access on 15 September 2026. Its announcement describes a trained model, a specialised architecture and a parallel sampling system. The training method is called Reinforcement Learning for Calibrated Decisions, or RLCD.

Jev reads supplied text and evaluates questions whose possible answers you define. TypeSafe calls this category “System One models”. Jev does not generate a written reply, code or a reasoning explanation. Its documented outputs are Choice, Score and Noul: selecting an option, evaluating ordered levels, or estimating the probability that a statement is true.

The public material does not establish Jev’s exact architecture, parameter count or starting model. I would not describe it as trained from scratch.

The training primer explains the intended calibration: across comparable predictions assigned a probability of 0.8, the outcome should occur about 80% of the time. That is useful for deciding when software should proceed or request review. It is not a guarantee about an individual answer.

How Is This Different from AI We Already Have?

My first comparison was with a language model returning JSON. If I can already ask an LLM to choose a department, what am I gaining?

Classification and probabilistic outputs are established capabilities. A classifier can return a label and scores without writing prose. Zero-shot classifiers can also evaluate labels supplied at runtime, including labels not encountered during training.

General-purpose LLMs can return constrained outputs too. OpenAI’s Structured Outputs enforces supported schemas, although the values can still be wrong. Valid JSON is not the novelty.

The comparison is about useful decisions, calibrated probabilities, latency and cost. TypeSafe’s headline speed gains come from particular vendor evaluations, not comparisons with every classifier or efficiently configured LLM.

The announcement’s “cannot hallucinate” claim concerns constrained outputs. Jev can still select the wrong permitted answer. A valid “billing” result can send a technical problem to the wrong queue.

Where Does OpenAI’s Decisions API Fit?

The names describe different layers. Jev is TypeSafe’s model, accessed through its System One API. OpenAI’s Decisions API is a dedicated service powered by GPT-6 Luna. The overlap in what developers receive is substantial.

CapabilityTypeSafe JevOpenAI Decisions API
Estimate whether a statement is trueNoulPredicate
Select from defined optionsChoiceChoice
Evaluate ordered levelsScoreScore
InputText, including JSONText and images
Endpoint/v1/systemone/v1/decisions

The Jev specifications and official OpenAI Decisions guide document these interfaces. OpenAI’s guide does not disclose an equivalent to TypeSafe’s RLCD training recipe. Similar outputs do not establish identical internals, and the documentation does not justify calling OpenAI’s offering merely a JSON wrapper.

OpenAI’s changelog records the public beta on 6 October. It reports decisions about ten times faster than the Responses API. That comparison is not a head-to-head result against Jev.

The trend extends beyond these two companies. Cloudflare released Clef and Clef-flash on 1 October, describing decision models compatible with Jev’s API. Several providers are making this component available. Which one belongs in an application still depends on its behaviour on that application’s work.

Context Comes First

The most useful question I asked was also the most ordinary: how can Jev answer anything meaningful if the information does not exist in the background?

Suppose a customer writes, “You charged me twice. Please return the extra payment.” The message provides evidence of a refund request. It does not prove that two payments occurred, that both settled, or that the customer is eligible for a refund.

A useful application must gather the relevant payment records and policy. The model brings language understanding and knowledge learned during training. It cannot know this account’s current transaction history unless the system supplies it.

Context + question + criteria → probabilistic judgement → application action.

I find that sequence more useful than thinking of Jev as an agent that takes over a job. Each part has an owner. Someone determines which records are relevant, defines the question, evaluates the judgement and implements the action.

More context is not automatically better context. An obsolete policy, a missing qualification or an unrelated customer record can mislead the model. Source status and access permissions should be handled by the application before evidence is passed along.

This builds on the retrieval problem in RAG Beyond the Demo: finding related text and finding evidence that supports an answer are different steps. A decision model gives us another way to evaluate that evidence. It does not remove the need to retrieve it well.

A Decision Model Needs a System Around It

This is where I had to correct the way I was describing the benefit. “A Scrum master can use Jev to flag incomplete tickets” leaves out the integration that makes that possible.

A backlog application could collect a ticket and the team’s agreed readiness criteria, ask Jev whether important information is missing, and display the result for review. Jev does not independently open the backlog, discover the team’s agreements or update the ticket.

In a playground, we supply the context manually. Paste a ticket, define a question and inspect the answer. No surrounding backlog system is required for that experiment, but no ticket is changed either.

Useful automation needs connections to data, application rules, a place to send results and a fallback when the model is uncertain or unavailable. It can be a small tool accepting pasted text or a file export. It can also be an integration within an existing product.

The model supplies a judgement. The application owns its consequences. Checking that acceptance criteria are present might flag a ticket for discussion. It should not quietly become an instruction to approve the ticket or judge the team.

Where Jev Could Fit Inside an AI Agent

To make the design concrete, I used a four-subquestion RAG workflow as a thought experiment. Imagine asking a product assistant:

Can our European team use the free plan for its monthly reporting workflow?

An LLM could decompose this into four questions: whether the service is available in the relevant region, whether reports can be exported on the free plan, what usage limits apply, and whether the required collaboration features are included.

Four subagents then search the relevant documentation. Jev would have a possible role after retrieval, before each agent turns its passages into an answer.

User question
  → LLM proposes four subquestions
  → four subagents retrieve evidence
  → decision model evaluates the evidence
  → application continues, retries or requests review
  → LLM combines supported findings

For the export subquestion, I would start with narrow checks:

  • Does this passage explicitly address exports on the free plan?
  • Does it support “available”, “unavailable”, or “insufficient evidence”?
  • Do the selected passages contradict each other about that capability?

The first two checks can evaluate individual passages. The contradiction check needs the relevant passages together. The model can only compare evidence included in its input.

Application code decides what follows. Sufficient evidence goes to the answering model. Missing evidence triggers a revised query, subject to a retry limit. Conflicting evidence can trigger retrieval of source metadata or review. A failed decision call needs an explicit fallback.

The LLM still generates subquestions, rewrites queries and writes the final answer. Jev evaluates defined questions. TypeSafe’s RAG passage-classification cookbook demonstrates this general placement between retrieval and generation; it does not establish results for my proposed four-agent workflow.

The smallest experiment I would run is to replace one existing evidence check. I would keep retrieval and generation fixed, then measure whether the replacement changes wrong acceptances, unnecessary retries, latency and total task cost.

What Could This Change for People and Teams?

For someone contacting customer support, the benefit would be fewer transfers and less repetition. A model could help recognise the issue and route it, but the support application still needs account context and working queues. That is a potential user benefit, not evidence that Jev has already reduced support delays at scale.

The same distinction makes the professional examples clearer:

RoleRelatable experimentWhat remains a human or system responsibility
Product managerClassify feedback into agreed themes and flag reports of blocked tasks.Validate the evidence and decide product priorities.
Agile coachCategorise recorded retrospective observations to surface recurring dependencies or unclear requirements.Explore causes with the team and choose improvements.
Scrum masterFlag work items that appear to lack acceptance criteria or contain unresolved dependencies.Facilitate clarification and help remove impediments.
Software engineerEvaluate retrieved passages or identify an agent repeating the same failed approach.Design the workflow, fallback and verification.

These are proposed applications. They require relevant records and an integration, and their usefulness needs testing. A model’s score is not a measure of team performance.

There is already a reported engineering implementation. Trau describes using Jev for ticket complexity, model routing, readiness signals, duplicate hints and repair-stall detection. Its readiness signals inform a human rather than block work. The authors also state that they have not yet benchmarked accuracy on their own data.

That is evidence of how the component can be integrated, not proof that it improves delivery. It connects with the handover problem in my article on AI and Agile: passing information faster only helps when the next person receives what they need to act.

How to Try Decision AI Today

I would begin with the export question from the RAG example. It is small enough to inspect without building the full agent.

TypeSafe’s quick start links to its Playground. OpenAI’s guide links to the Decisions Playground. Sign in and configure equivalent Choice questions, using the same instructions and answer options in both.

Ask, “What does this passage establish about exports on the free plan?” Define three options: “available”, “unavailable” and “insufficient evidence”. Then try these fictional passages separately:

PassageReference answer
Report exports are available on paid plans only.Unavailable
The free plan includes dashboards and saved reports.Insufficient evidence
All plans support CSV report exports.Available

The second passage is the important one. It mentions reports but does not establish whether they can be exported. I would want an evidence check to preserve that gap.

Next, supply the first and third passages together and ask whether they conflict about free-plan exports. Then add source dates and an explicit rule about which policy is current. Observe whether the result follows the supplied rule instead of merely favouring confident wording.

Record reference answers before running either model. These cases form a tiny gold-standard evaluation set. They are different from the questions sent to Jev during execution.

Expand the set with paraphrases, exceptions and missing evidence. Keep some cases aside while adjusting instructions. Compare accuracy and time, and inspect confidently wrong answers. A higher reported confidence is not a win.

TypeSafe’s confidence documentation explains that its separate confidence field is derived from the probability distribution. It is not an independent verification of the answer. I would not assume the same numerical threshold has the same meaning across different services.

What Would Make This Shift Matter?

I began by asking what Jev had built and how OpenAI’s offering compared. The RAG example led me towards a more useful decision: which small judgement in a workflow is worth evaluating separately?

It could be whether a passage supports an answer, whether a ticket needs clarification, or which existing handler should receive a request. If that judgement is frequent, a dedicated decision service might change the economics of checking it. It might also add a dependency without improving the outcome.

I would test the complete workflow. Does it answer more questions correctly? Does it stop when evidence is missing? Does it avoid unnecessary retrieval rounds? What happens when the service fails? Are the extra calls justified by better behaviour?

The shift becomes meaningful when developers can place useful, measurable judgements throughout software and people experience less waiting, repeated explanation and avoidable rework. The context, integrations and policy remain part of the work.

Before adding a decision model, identify the evidence it needs, the judgement it must make and the action its answer will trigger. Then test that chain.

Sources

Sources checked on 7 October 2026. The RAG and role-based experiments above are proposals, not measured results.

  • OpenAI: Decisions API video. The user-supplied starting point for this exploration. Product details rely on the documentation below, not a verified transcript or reproduction of the demonstration.
  • TypeSafe: team. Supplies founder names, roles and stated professional backgrounds. Company biographies are first-party accounts.
  • Ouyang and colleagues: Training language models to follow instructions with human feedback, 2022. Lists Almeida as a co-author and describes InstructGPT training. It supports his research contribution without implying that one person invented RLHF.
  • TypeSafe: Introducing System One Models & Jev. Supplies launch and technical claims. Vendor comparisons are workload-specific; constrained outputs can still be incorrect.
  • TypeSafe: System One. Defines Jev’s three primitives and the absence of generated replies, code and explanations. It is an interface description, not an independent quality evaluation.
  • TypeSafe: AI primer. Describes RLCD as a post-training approach and explains calibration across groups of predictions. It does not disclose a complete reproducible training recipe.
  • Hugging Face: zero-shot classification. Documents classification against runtime-supplied labels, showing that flexible label spaces predate these decision services. Classification scores alone do not establish calibration.
  • OpenAI: Structured Outputs. Documents schema adherence and the possibility of semantic mistakes. It establishes an existing comparison for structured answers.
  • TypeSafe: models. Documents text input and the System One endpoint. Published model details do not establish the undisclosed architecture or starting weights.
  • OpenAI: Decisions guide. Documents GPT-6 Luna, question types, image support, Playground access and the Responses API speed comparison. It does not establish internal equivalence with Jev or a head-to-head advantage.
  • OpenAI: API changelog. Records the 6 October 2026 beta release. Public release timing does not establish when internal development began.
  • Cloudflare: introducing Clef. Establishes the 1 October release of Clef and Clef-flash and claimed Jev API compatibility. This supports the wider trend, not a universal performance ranking.
  • TypeSafe: classifying RAG passages. Demonstrates checks between retrieval and generation. Its example is not a benchmark for the proposed four-subagent architecture.
  • Trau: small decisions in an autonomous coding pipeline. Reports ticket assessment, routing and loop checks using Jev. The authors explicitly say they have no accuracy benchmark on their own data yet.
  • TypeSafe: quick start. Documents manually supplying state and questions in the Playground. A Playground judgement does not execute an application workflow.
  • OpenAI: Decisions Playground. The signed-in experiment interface linked from the official guide. This article does not report an authenticated trial.
  • TypeSafe: confidence. Explains the probability-derived confidence statistic and domain-specific threshold selection. Confidence is not an independent correctness check.
#decision-ai#jev#openai#rag#context-engineering#ai-agents

Share this article

Pass it on to someone who might find it useful.

LinkedInFacebookXWhatsAppEmail

From theory to practice

See how the ideas become working systems.

I’m working on hands-on examples and videos that build real agentic workflows step by step. Join the list for new articles, practical material, and the first course updates when they’re ready.

Practical AI notes, no noise. Unsubscribe whenever you like.

Discussion

What did this make you think about?

Share what you have seen in practice, ask a question, or add a different perspective.

Keep exploring

More in Working with AI

  1. When the AI Bubble Bursts, What Will Still Be Worth Learning?24 September 2026
  2. Build an AI Agent from First Principles Before You Choose a Framework17 September 2026
  3. OpenAI’s Agents API: What It Is and What Changes for Developers17 September 2026
View all in Working with AI→

Explore other areas

Using AI→AI Fatigue Is Real and Workplace Jargon Is Making It Worse
Building AI→What Building AI Means: My First Step as a Full-Stack Developer
Browse the complete archive→
Codehance emblemCODEHANCE

An open notebook from an AI Lead at Gravity9 on using AI, working with AI, and building AI.

The three layers

  • Using AI
  • Working with AI
  • Building AI

This blog

  • Latest notes
  • Complete archive
  • RSS feed
  • support@codehance.com

© 2026Codehance Ltd. All rights reserved. Registered in England & Wales. blog.codehance.com