Jev vs LLMs in 2026: When Should AI Decide, Not Write?
I used Claude, an AI assistant, to help research, draft and edit this post and to build its diagrams. I ran the tests myself, and every number traces to my raw results.
TypeSafe AI says its new model, Jev, is 193x faster and 445x cheaper than rival models. Those figures come from TypeSafe's own tests, against Claude Sonnet 5 for speed and Claude Opus 5 for cost.
What is Jev? Jev is an AI model from TypeSafe AI that picks answers from a list you give it, instead of writing them. Ask it about a support ticket (which team, how urgent, how angry?) and it returns three answers, with a probability for every option.
In September 2026 I ran Jev and fifteen large language models (LLMs), the kind of AI behind ChatGPT, on the same questions. The bottom line: faster, but nowhere near 193x.
This is a decision brief for customer experience (CX) leaders. It gives the verdict first, then how Jev works, how far to trust it, where it fits, and a test to run before you automate anything.
"In my test" below points to the companion post, Is Jev Really 193x Faster and 445x Cheaper? (TK: companion post URL), which holds the full method and every result.
Is Jev really 193x faster and 445x cheaper in 2026?
Not in my test, and not in any independent test I found. Jev was the fastest model I tested and the cheapest on short tickets, but only about as accurate as most small LLMs. Its edge was single or double-digit multiples, not hundreds.
In my test, against the fifteen LLMs (eight everyday models and the makers' seven priciest):
Speed: Jev was the fastest model in every test, at 0.38 seconds per ticket. That is 2 to 6x faster than the everyday models, and about 6x faster than Claude Sonnet 5, the rival behind TypeSafe's 193x.
Cost: Jev was cheapest on short tickets, at $66 a year for 10,000 tickets a day. GPT-6 Luna, the cheapest LLM, came to $165 (both scaled up from three tickets).
Accuracy: Jev scored 81% on 200 banking questions, level with six of the eight everyday models and behind the other two. The seven priciest models scored 90 to 96% on a 50-question subset, against Jev's 80%.
As a first filter: Jev answered what it was 90% or more sure of and passed on the rest. Before a pricey model, this kept accuracy at about a third of the cost. Before a cheap model, it added cost for little or no gain.
I can't rule out bigger gains on TypeSafe's long, multi-step jobs, which I didn't test.
My assessment: Jev's case rests on speed and price. On accuracy it was level with most of the everyday models and behind the priciest ones.
How did I test Jev against 15 LLMs?
The models: Jev 1.13 and fifteen LLMs from Anthropic, OpenAI, Google, xAI, Moonshot and Alibaba. Eight were everyday models, the small and mid-sized ones teams use for sorting. Seven were the makers' priciest.
The questions: three support tickets, and up to 200 messages from Banking77, a public collection of short messages to an online bank. Each must be sorted into one of 77 topics, such as a late card or a blocked PIN.
The access: I reached Jev through OpenRouter, a service that offers many AI models through one account, since I couldn't get direct access from TypeSafe at the time.
The limit: I didn't test TypeSafe's long, multi-step jobs, which is where its 193x figure comes from.
What did TypeSafe's own chart and independent tests find?
TypeSafe's headline doesn't name its rivals, but its own chart shows each number uses a different one. The 193x is against Claude Sonnet 5; the 445x is against Claude Opus 5, the most expensive model in its test.
On TypeSafe's own chart, against GPT-5.6 Terra, which scored virtually the same as Jev, the gap was about 25x faster and 77x cheaper. TypeSafe itself calls its multipliers "on the higher end of real world gains", from four long jobs its own team wrote.
Testers published their own checks within a week of launch. These are small blog tests, none peer reviewed. ("Calibrated" below means a model's confidence matches how often it is right.)
AY Automate, 791 labelled decisions: 2.0 to 3.5x faster than small LLMs, with accuracy level with them. "The 193.6x and 444.6x figures ... did not show up."
Aman Kumar, about 16,000 calls: level with or ahead of two small OpenAI models on 3 of 4 datasets, at 5 to 56x lower cost. Worse on whole documents.
WotAI, 150 passages: the best calibrated of the models that answer in under a second, a hair ahead of Claude Haiku 4.5. Claude Sonnet 5 was better calibrated still, but 3.7x slower.
So the headline multiples did not hold up. What is left is a fast, cheap model that works differently from an LLM.
What is Jev, and what is a "System One" model?
Jev takes a piece of text and a list of questions with fixed answers, and returns an answer to each with a probability attached.
TypeSafe AI came out of stealth on September 15, 2026, with Jev as its first model and a $40M seed round led by DCVC. Its CEO, Diogo Almeida, spent about four years at OpenAI and is a co-author of the InstructGPT paper.
The psychologist Daniel Kahneman described two kinds of thinking in Thinking, Fast and Slow: fast, automatic System 1 and slow, deliberate System 2.
Routing, tagging and flagging are System 1 work, and TypeSafe calls Jev a "System One" model. LLMs are System 2 machines, and since 2024 reasoning models, which "spend more time thinking before they respond", have pushed them further that way.
We have been using essay-writing AI for System 1 jobs because it was the easiest AI to reach.
Jev is not the only model built for one narrow job. On the System 2 side, Salesforce's Koa is a CRM-specific reasoning model.
Every Jev request has two parts. The state is the text you want judged: a ticket, a lead record, a product review. The questions are what you want to know, and each is one of three types:
Choice: pick one option from a list you define, such as billing, technical or sales.
Score: a place on a scale you define, such as calm, frustrated or very angry.
Noul: TypeSafe's name for a yes-or-no question, such as "is this urgent?"
How did business AI get from writing to deciding?
LLMs were designed to write. Since 2020, businesses have bent them into making quick decisions.
An LLM learns one skill: guessing the next word. OpenAI's GPT-2 was "trained simply to predict the next word in 40GB of Internet text" (OpenAI, 2019). So whatever you ask, an LLM answers by writing.
In 2018 each job got its own small trained model, such as Google's BERT, taught on thousands of past tickets with the right answer attached.
In 2020 OpenAI opened GPT-3 to other companies' software as a "text in, text out" service, and ChatGPT launched on November 30, 2022.
In 2023 OpenAI let developers request replies as data and added JSON mode, a labelled format that software reads. The reply was always readable, but not always with the labels you asked for.
In 2024 Structured Outputs forced every reply into an exact template. On OpenAI's test, an older GPT-4 filled a complex template correctly under 40% of the time; the new model, every time.
Structured outputs fixed the shape of the answer, not the cost or the doubt. Here is one support ticket answered four ways:
The LLM still writes its answer out piece by piece. You pay for every piece and wait while it writes. And "technical" looks the same whether the model was 99% sure or flipping a coin.
What does a Jev answer look like next to an LLM's?
On the same support ticket, Jev and an LLM gave the same three answers, but only Jev showed how close each call was. The ticket is the sample from TypeSafe's quickstart:
"Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP."
I asked Jev three questions about it: which team should handle it (a Choice), how frustrated is the customer (a Score) and is it urgent (a Noul). Here is its answer:
I asked Claude Haiku 4.5, a small, fast LLM, the same three questions, and also how sure it was of the team. Its reply had to fit an exact template.
Claude: technical, frustrated, urgent, and 0.95 sure of the team. Jev: technical 81%, billing 19%, sales 0%; frustrated; a 0.99 chance it is urgent.
Claude's 0.95 is text it wrote because I asked for it. Jev's 81% technical and 19% billing are the scores it used to pick its answer.
The difference showed on a second, deliberately ambiguous ticket in my test. Jev said billing was only about 70% likely and urgency about 62%: a clear "check this one". On urgency, the fifteen LLMs I tested split nine to six between yes and no.
A model that says "check this one" is only useful if the warning is reliable. That is the next question.
Can you trust Jev's confidence score?
In my test, a low confidence score flagged most of Jev's mistakes, and its most confident answers were right 94% of the time. A stricter independent test found Jev more confident than it was accurate. Treat the score as a claim to check on your own data.
The idea is simple: act on the answers Jev is sure of, and have a person check the rest. In my test, the score split Jev's answers cleanly:
On 200 banking customer questions, Jev was right only 3 times in 27 when under 70% sure. Between 70 and 90% it was right 21 times in 26, and at 90% or more, 138 times in 147 (94%).
These bands come from one test, so treat the actions in the diagram as a starting point, not advice.
Aman Kumar's tests found the same at the top: Jev's answers at 0.9 confidence or higher were right 90 to 100% of the time.
The percentages themselves weren't exact, though. In my test, most LLMs' written confidence numbers matched how often they were right at least as well as Jev's. The LLMs bunched on round figures such as 0.9 and 0.95, while Jev used 48 different values.
A stricter test was harsher. On 273 decisions, over half of them the tester's own internal routing questions, NXTG.AI found Jev 66% confident on average but only 53% right.
NXTG.AI judged Jev "not calibrated as claimed" on 4 of its 5 question types, meaning its confidence did not match how often it was right. For one of them, sorting contract clauses into 20 kinds, its advice was to use Jev's answer but not to automate on its confidence score.
For a Choice question, Jev's confidence is stricter than its probability: the Stripe ticket's 81% became a confidence of 0.72. TypeSafe's docs explain why: the confidence gives no credit for what a random guess would get right, one time in three with three teams.
How does TypeSafe say it makes those numbers honest?
TypeSafe's case is that ordinary LLM training never rewards a model for knowing how sure to be. LLMs are rewarded for answers people rate highly, and for answers a computer can check, such as maths and code.
People tend to rate confident answers higher, so a model can learn to sound sure whether or not it is.
There is evidence that this happens. OpenAI's GPT-4 report (Figure 8) found the raw model's confidence matched how often it was right on multiple-choice questions. After the training that made it a better assistant, that match got noticeably worse.
TypeSafe says it trains Jev more like a weather forecaster. A forecaster is judged on whether "70% chance of rain" days see rain about 70% of the time. A model that passes that test is called calibrated. TypeSafe calls its method Reinforcement Learning for Calibrated Decisions.
Jev learns only from computer-generated practice cases, TypeSafe says, so the right answer is always known in advance.
TypeSafe hasn't published the details: no scoring rule, no account of how the practice cases are made, no calibration charts and no paper. So treat "calibrated" as a claim to check on your own data, not a fact. Even an honest 0.9 is wrong about 1 time in 10.
Can't I just use an LLM with a JSON schema?
For the shape of the answer, an LLM with a JSON schema (a fixed answer template) does as well as Jev. OpenAI's Structured Outputs and Claude's equivalent block any reply that breaks your template, including a team that is not on your list.
Three differences remain:
Odds for every option. An LLM returns one answer. Getting its odds for the alternatives is awkward: the Claude API doesn't offer them, and OpenAI withholds them from its reasoning models.
Speed. An LLM writes its answer piece by piece. TypeSafe says Jev scores all its questions in one pass, though it hasn't published the design.
Training. LLMs are trained to be helpful and correct; Jev, TypeSafe says, to be calibrated.
LLMs still win at a lot. They can reason before answering, handle long or messy documents, explain why they chose an answer, and write.
Jev's speed trick can be partly copied, critics add. Engineer Sean Goedecke got a 2 to 3x speedup from an ordinary small model that anyone can download and run. He suspects Jev has no "substantial technical moat."
Where does Jev fit, and who should not use it?
Jev fits short, high-volume decisions with fixed answers, such as routing, tagging and flagging support tickets. By TypeSafe's own account it is weaker on large inputs, languages other than English, maths and dates. Test it on your own past tickets before you automate anything.
Jev is worth testing when all four of these are true:
You can list the possible answers in advance: teams, categories, yes or no, a scale.
The input is short and focused: a ticket or a record, not a 40-page contract. TypeSafe lists large input as a weak spot, and one independent test found Jev worse on whole documents.
The volume is high, thousands of decisions a day or more, so speed and cost add up.
A wrong answer is cheap to catch or undo, or you can send low-confidence cases to a person.
In a support or customer relationship management (CRM) process, that usually means Jev sits around the LLM rather than replacing it:
Jev makes the calls and the LLM does the writing, so you pay LLM prices only for the step that needs language. The hand-off to a person needs its own design; in Salesforce, for example, escalation in Agentforce runs through a dedicated flow.
Keep firm rules, such as refund limits and who may approve what, in code rather than in the model, as TypeSafe's build guide advises. Agentforce teams face the same choice of when to script an agent's rules.
Who should not use Jev? Anyone whose job runs into these limits:
Size: TypeSafe allows 64,000 tokens (word pieces) per request, with 32,000 for the text being judged and the longest question. By the usual rule of thumb, that is about 24,000 words or 50 pages. A 40-page contract fits, but TypeSafe says accuracy shifts as the text grows.
Text and language: Jev reads text only, and TypeSafe's docs say languages other than English "are handled but not equally well".
Control: you can't train it further on your own data, and it runs only on TypeSafe's servers.
Weak spots: TypeSafe lists nine in all, including maths, dates and a lean toward the option listed first.
And of its price, TypeSafe says: "We can't prove it isn't subsidized."
How do you test Jev on 500 of your own tickets?
Jev's confidence is only worth routing on once it has held up on your own data. To test it:
Take 500 past tickets where you already know the right answer.
Run them through Jev and record each answer and its confidence.
Group the results: 0.9 and above, 0.7 to 0.9, and below 0.7.
In each group, count how often Jev was right.
Aim for at least 100 tickets in each group, adding past tickets if you need to. If the top group is right at least 90% of the time, you can automate that band with evidence behind it. If it's right only 70% of the time, you've avoided automating a problem.
Then set your own cut-offs by what's at stake. A wrong tag is cheap to fix, so you might automate tagging at a lower score than my diagram suggests. A refund might need 0.95 and a person anyway.
Pin the version you tested (jev-1.13.0, not jev-latest), so an update can't quietly shift your numbers.
Should you use Jev in 2026? The decision on one screen
Jev earns a trial only when four conditions hold, and a place in your process only after it passes a test on your own data. Here is the whole decision:
Jev is named after William Stanley Jevons, the 19th-century economist (TypeSafe's launch post says so). Jevons noticed that when steam engines burned coal more efficiently, Britain burned more coal, not less.
TypeSafe is betting the same will happen with AI decisions. When a judgement call costs a fraction of a cent, you run one on every ticket, every lead and every agent action.
Whether Jev itself wins is an open question. My assessment: the idea behind it holds either way. Use System 2 models for writing and reasoning, and System 1 models for the thousands of quick decisions a business makes every day.
Match the model to the kind of thinking the job needs.
Want a second opinion on your own 500-ticket test? Book an architecture consult
Related Readings
Let’s Talk
Drop us a note, we’re happy to take the conversation forward 👇🏻

