Is JEV an AI agent? Inside TypeSafe's "System One" models
Sent a support ticket, TypeSafe's Jev doesn't write a reply. It returns a number: 99.9% probability the message is urgent. That trade — give up generated text for a typed answer in milliseconds — is the whole idea behind a "System One" model, and it is not the same thing as an AI agent.
Jev is not an AI agent. It's a "System One" decision model: sent a fixed list of possible answers, it returns one of them with a calibrated probability, in one fast pass, and stops there — no conversation, no tool calls, no multi-step goal. TypeSafe reports it as 40–200x faster and up to 444.6x cheaper than an LLM call for that kind of decision; independent benchmarks confirm large but more modest gains. In practice, teams are dropping it inside agent pipelines as a routing and guardrail layer, not using it to replace the agent itself.
Diogo Almeida spent four years chasing a question. At OpenAI, he helped build the instruction-following research that became ChatGPT [1] — and watched language models turn superhuman at chat while automation, the thing businesses actually wanted, stayed stubbornly hard. On 15 September 2026 his startup, TypeSafe AI, shipped what he calls the missing half: a "System One" model, named for Daniel Kahneman's fast, intuitive System 1 thinking [1]. Its first public release is Jev — named after William Stanley Jevons, on the bet that cheaper intelligence creates more demand for it, not less [1].
Jev does not generate text. Sent a "state" — a paragraph, a support ticket, a game board — plus a set of typed questions, it returns typed answers with calibrated probabilities, in one parallel pass rather than one token at a time [1][2]. TypeSafe's docs define three question types: Choice picks from a fixed list ("which team should handle this ticket?"), Score rates something on an ordered scale ("how frustrated is this customer?"), and Noul answers yes-or-no with a probability ("does this message request a refund?") [2]. Every answer's shape is fixed in advance — in TypeSafe's own words, the result is "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out" [1].
A model that refuses to write a sentence turns out to be exactly what a routing decision needed.
The numbers, checked twice over
TypeSafe's own published comparison puts Jev's input tokens at $0.042 per million — $42 per billion — against $0.20 to $10 for the LLMs it benchmarks against, with output free because there isn't any text to bill for [1]. On speed, its workflow evaluations claim 70–500 millisecond responses against 3 to 329 seconds end-to-end for frontier LLMs, and the "193.6x faster, 444.6x cheaper" figure on its homepage traces to one logged run: $0.000081 and 0.114 seconds for Jev against $0.013880 and 8.566 seconds for the comparison model [7]. TypeSafe flags the bias itself — its reference answers average two OpenAI and Anthropic models, and it says plainly its published figures are "on the higher end of real world gains" [1].
Two outside checks, both published within a week of launch, land closer to earth. LiteLLM's engineering team ran Jev against Claude Haiku 4.5 as a request classifier inside their own AI gateway: across 240 calls, Jev's median classification time was 126.81ms against Haiku's 688.40ms — 5.43x faster — at 96.1% lower registry-priced cost, and it matched the team's own routing labels on 95% of calls against Haiku's 73.75% [4]. LiteLLM is upfront about the ceiling on that number: the labels were authored by the same team without independent review, on a corpus of 80 synthetic prompts, and the result "does not establish general classification accuracy or the quality of the final answers" [4].
A second, scrappier check came from an independent developer, Akshay Kanthed, who built a tool that scans a codebase for LLM calls that are secretly just routing or classification, then benchmarks converting them [5]. Against the live OpenAI API on three real test cases, a hand-written local rule beat a GPT-4o-mini call by 1,240x to 12,483x in latency, at effectively zero cost [5]. He's explicit that the comparison is illustrative, not definitive: the "local rule" only stands in for Jev to show the shape of the saving, "not a guarantee that your specific decision generalizes as cleanly" [5].
Where it actually plugs in
None of this makes Jev a replacement for an agent — TypeSafe doesn't claim that, and neither does anyone benchmarking it. LangChain's own writeup is blunt: "Jev isn't a drop-in replacement for an LLM… That makes it a promising complement to the model driving your agent: use an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way" [6]. Their integration ships two concrete patterns: a model-routing middleware that has Jev pick a cheap or a powerful model per request before the agent starts, and an "auto mode" middleware that has Jev screen a risky tool call — a bash command, say — before it's allowed to run [6].
A research paper posted the same week goes further, using a System One model as the control layer for an agent's entire memory system rather than one decision. Jev-Mem splits the work into a fast "System One" plane that handles memory typing, retrieval routing and when to stop searching, and a slower "System Two" plane — an LLM — called only for the final reasoning and answer [3]. On the LoCoMo benchmark, that split cut memory construction time to 158 seconds — a 6.6x speedup over the fastest competing system it tested against — and average query latency to 0.93 seconds, a 36.7% reduction, while its answer-quality score rose 11.0% over the strongest baseline [3]. The pattern is the same across all three examples: Jev sits in front of or beside the reasoning model, deciding what happens next. It doesn't do the reasoning itself.
What "zero hallucinations" leaves out
TypeSafe's boldest claim is that Jev can't hallucinate, and on one narrow definition that's true: because every answer has to match a schema declared in advance, a type error is "mathematically impossible" rather than empirically rare [1]. TypeSafe's own docs are more careful about what that guarantees: calibration is "measured across groups of predictions; it does not guarantee that an individual answer is correct" [2]. A confidently wrong typed answer — the right shape, the wrong value — is still on the table. What Jev removes is the chance of a malformed response breaking your code downstream, not the chance of a wrong decision.
So — is it an AI agent?
By the definition on this site, no. An AI agent decides its own next step across a sequence of actions; Jev answers one typed question and stops. It doesn't hold a conversation, call tools on its own, or pursue a goal over multiple turns. It's a component other systems — agents included — can call when they need a fast, cheap, structured answer rather than a written one. Whether that's more useful to you than the LLM already inside your agent depends on how much of what that agent does is really a choice between a short list of options — which, per the three benchmarks above, is more of it than most teams have measured.

Sources
- TypeSafe AI — "Introducing System One Models & Jev" (blog, 15 Sep 2026)
- TypeSafe AI — Docs: "System One" (docs.typesafe.ai/concepts/system-one)
- Jiang, Li & Li — "Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents," arXiv:2609.23986 (21 Sep 2026)
- liteLLM — "JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost" (benchmark, 20 Sep 2026)
- Akshay Kanthed, DEV Community — "I stopped believing '99% cost reduction' claims, so I benchmarked my own tool instead" (21 Sep 2026)
- LangChain — "Building a Harness with Jev" (17 Sep 2026)
- TypeSafe AI — homepage (typesafe.ai), workflow cost/latency figures
Frequently asked questions
Is Jev an AI agent?
No. Jev is a "System One" decision model: it answers one typed question at a time — a choice, a score, or a yes/no — from a fixed set of possible answers, and stops. It doesn't hold a conversation, call tools on its own, or pursue a multi-step goal, which is what defines an agent on this site. It's typically used as a component inside an agent's pipeline, not as a replacement for the agent.
Is Jev really hallucination-free?
It's guaranteed to never return an answer outside its declared schema — TypeSafe calls this mathematically impossible rather than empirically rare. That's different from being correct: TypeSafe's own documentation notes that a confident, correctly-typed answer can still be wrong, since calibration is measured across groups of predictions, not guaranteed for any single one.
How much faster and cheaper is it than a normal LLM call?
TypeSafe's own workflow evaluations claim roughly 40–200x faster responses at similar intelligence for this kind of task, at up to 444.6x lower cost. Independent benchmarks land lower but still large: liteLLM measured Jev's median classification latency at 5.43x faster than Claude Haiku 4.5 with about 96% lower cost, on an 80-prompt test the authors describe as narrow and self-authored rather than independently reviewed.