ExecutionNew·Falk Gottlob··7 min read

Jev Is 400x Cheaper Per Call. Price the Wrong Pick.

Three creators explained Jev by price per call this week. My classifier went from 89% to 61% in six weeks. Rank branches by what a wrong pick costs to undo.

ExecutionJevdecision modelsTypeSafemodel costagent evalsreversibilityconfidence calibrationBasia KubickaClaire VoJohn LindquistWinston HuynhLangChainVinay GoelRam SomaAmplitudekill list
Helpful?

Execution Falkster cover on teal: a railway switch lever standing where one track splits into two, with a blank paper price tag tied to the lever.

Three people I read explained Jev this week, and all three led with what a call costs. The cost is real and the model looks useful. But price per call is the wrong test for moving a step in your agent onto it, and I have a failed launch that says why.

The short version

Jev, TypeSafe's decision model, picks from answers you supply and returns a probability. Basia Kubicka's October 6 explainer puts it at $0.042 per million input tokens with free output, Claire Vo and John Lindquist showed eight demos on How I AI, and Winston Huynh at LangChain reported a full set of eval judgments at $0.34 on Jev against $28.17 on a frontier model. Kill price per call as the comparison. Vinay Goel and Ram Soma at Amplitude cite a pilot with 87.4% mean confidence against 75.3% accuracy, and note that escalating uncertain answers does not catch confident wrong ones. A sentiment classifier of mine scored 89% on held-out data and 61% in production within six weeks. A cheap call does not make a wrong pick cheaper. It puts a pick at every branch. Rank the branches by what a wrong pick costs to undo, and move the cheap-to-reverse ones first.

Three posts, one number

I read Kubicka's post in full, on LinkedIn. It is the clearest explanation of the thing I have seen. An LLM answers by writing, one word at a time, and bills for each. Jev answers by picking. Ask both whether an email is urgent: the LLM writes a sentence with a reason, and Jev returns "yes, 91%." Her arithmetic: LLM input runs $0.20 to $10 per million tokens with output about five times more, Jev input is $0.042 per million with output free, so a million calls of 500 tokens is about $21 on Jev and $100 to $5,000 on an LLM before a word of output. TypeSafe's own tests, she reports, put it at up to 444x cheaper.

Claire Vo's How I AI episode with John Lindquist, on Lenny's feed, is titled for the fastest, cheapest model he has used. I read the show notes and did not watch the episode, so I will only say what the notes list: eight demos, from a voice to-do app to an app router, and a segment on where Jev falls short.

Winston Huynh's LangChain post puts Jev into LangSmith as an eval judge. In LangChain's test, Jev matched a human reviewer on every decision, and the full set of judgments cost $0.34 on Jev against $28.17 on Claude Sonnet 4.6. He is careful to add that this was one test on one agent.

None of them oversell it. Kubicka lists where the LLM still wins. Huynh says LLM judges are not obsolete. Still, the headline in all three is a price, and a price is what a team will paste into the pull request.

The classifier I shipped

A decision model takes text in and returns one of a few labels. I shipped one of those before anyone called it that. It is in 10 AI Agents I Built That Failed. The Honest Retrospective.: a sentiment classifier that scored 89% on a held-out set and 61% in production within six weeks.

It sat on a branch. Angry tickets went to senior support reps, faster. That is the first use on Kubicka's list: which team gets this ticket, and how urgent is it.

The held-out set was English-heavy. Customers writing in other languages were scored as angry no matter what they wrote. Nothing errored. Every answer came back in the right shape. Non-English tickets that did not need a senior rep went to one, and the English-speaking customers who were angry waited behind them.

Price per call never came up in that story. What it cost was response time on critical English-language tickets, up 18% for two months before I caught it.

What the accuracy posts say

Vinay Goel and Ram Soma at Amplitude published the other half on September 28, and I read it in full. They are deploying Jev themselves, so this is not a takedown. Three things from it:

A 77-case pilot they cite found 87.4% mean confidence against 75.3% actual accuracy. About 12 points of overconfidence.

In a 791-decision study, a cascade that escalated below 0.80 confidence matched frontier accuracy at about a quarter of the cost. But the decision model agreed with the frontier model 6 to 8 points more often than it matched the ground-truth labels. Five confident errors on near-identical intents were missed by both.

And they quote Anthony Maio at Pieces: the type system constrains the shape of the output, not the judgment. Clean types, high confidence, no failed check, and the ticket still went to the wrong queue.

That last line is my classifier.

What a cheap call changes

It does not change what a wrong pick costs. It changes how many picks you make.

At LLM prices, a team puts a model call at a few branches and a person or a rule at the rest. At $21 per million calls, there is no reason not to put a decision at every branch, and the Amplitude post says many teams are doing that. In The Cost of Being Wrong Is the Only Number That Matters Now I argued that an agent makes a wrong call confidently, consistently, and at volume, and that the triage is reversibility, not size. A decision model is that argument with the price removed as a brake.

The page to write first

Kubicka's close is the right step one: count how many of your agent's steps are choices. Then, for each choice, four lines.

  1. The branch, and what a wrong pick does downstream. Wrong queue, wrong tool, wrong customer message.
  2. Whether you can undo it on Tuesday. Tagging reviews by topic, yes. Gating whether a command is safe to run, no. Cheap-to-reverse branches move first.
  3. A labeled set for that branch, split by the slices you care about. Language, segment, product line. An average hides the slice, which is the point of Which Slice Holds the 3%?.
  4. Who re-runs it in week six. A name. My 89% was true on the day it was measured.

The comparison I would put in the pull request is then two numbers per branch: what the calls cost, and what last month's wrong picks cost at the measured error rate for the worst slice. On most reversible branches Jev will win that easily. On the others you will know what you are buying.

This sits in Enterprise AI Agents, where the thesis ends on what each successful outcome costs. A call is not an outcome.

Pick one thing this week. Take the step you most want to move to a decision model and write line two for it. If the answer is no, build the labeled set before you touch the code.

Related answer: Is price per call the right way to compare a decision model like Jev?

Sources: Basia Kubicka, "JEV, an AI model up to 400x cheaper than LLMs, launched 3 weeks ago," LinkedIn, October 6, 2026. Claire Vo with John Lindquist, "Jev: 8 real use cases for the fastest, cheapest model I've ever used," How I AI on Lenny's Newsletter, September 30, 2026 (show notes). Winston Huynh, "Jev is now available in LangSmith Evals," LangChain blog. Vinay Goel and Ram Soma, "An analysis of Jev: Faster and cheaper, but watch the accuracy," Amplitude Blog, September 28, 2026, which is also the source for the pilot, the 791-decision study, and Anthony Maio's remark. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026.

Share this post

Also on Medium

Full archive →

Frequently asked

What is Jev?+

Jev is TypeSafe AI's decision model, which TypeSafe calls a System One model. It does not generate text. You send it an input and a set of questions, each a choice from a list, a score on a scale, or a yes or no, and it returns a typed answer with a probability for each. Basia Kubicka's example: asked whether an email is urgent, an LLM writes a sentence and Jev returns yes, 91%.

How much cheaper is Jev than an LLM?+

Per Basia Kubicka's post of October 6, 2026, Jev input costs $0.042 per million tokens and output is free, so 1 million calls of 500 tokens is about $21, against $100 to $5,000 on an LLM before output. She reports TypeSafe's own tests at up to 444x cheaper. LangChain reports that in one test on one agent, a full set of eval judgments cost $0.34 on Jev against $28.17 on Claude Sonnet 4.6.

Why is price per call the wrong way to compare a decision model?+

Because it prices the call when the pick is right. The cost that matters sits in the wrong picks: what each one does downstream and how long it runs before anyone notices. A cheaper call also means a decision at every branch, so the number of picks goes up. A sentiment classifier I shipped scored 89% on held-out data and 61% in production within six weeks, and price per call was never the issue.

How accurate are decision models like Jev?+

It depends on the task and the data. Vinay Goel and Ram Soma at Amplitude cite a 77-case pilot with 87.4% mean confidence against 75.3% actual accuracy, a benchmark site that reports a range of 62.6% to 95.4% and declines to publish one headline number, and a 791-decision study in which the decision model agreed with a frontier model 6 to 8 points more often than it matched the ground-truth labels. In Amplitude's own synthetic benchmark Jev caught every meaningful chart and flagged more benign ones than the best general models.

Does escalating low-confidence answers to a bigger model fix accuracy?+

It catches uncertain answers. It does not catch confident wrong ones. The Amplitude analysis notes that in the 791-decision study, five confident errors on near-identical intents were also missed by the frontier model, which is the failure the cascade is supposed to catch.

Which agent steps should move to a decision model first?+

The ones where a wrong pick is cheap to undo. For each branch, write down what a wrong pick does downstream, whether you can reverse it on Tuesday, a labeled set for that branch split by the slices you care about, and who re-runs that set in week six. Branches that touch a customer, money, or a data model wait until the labeled set exists.

THE SHORT ANSWER

PART OF

AI Product Management

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.