AI AgentsNew·Falk Gottlob··6 min read

Label Before You Route

LangChain's model router cut median cost per thread 64%, proven by an A/B on its own engineers. For a customer-facing agent, label 150 real threads first.

AI Agentsmodel routingmodel routerharnessLangChainOpen SWESydney RunkleEugene Yurtsevevalscost per outcomeSmartcatagent drops
Helpful?

AI Agents Falkster cover on green: a railway switch lever with a paper luggage tag tied to its handle, standing where one track splits into diverging lines.

LangChain published the numbers on a model router for its coding agent on October 1, and they are good enough that a lot of teams will have a router on the roadmap by next week. This is the step I would put in front of it, and it takes an afternoon.

The short version

Sydney Runkle and Eugene Yurtsev at LangChain built a model router into the harness of Open SWE, their coding agent. It picks one of three model tiers on the first message of a thread. In an A/B test across 973 threads, the median routed thread cost $0.94 against $2.61 on the always-top-tier control, 64% less, with merged PRs flat at 29.2% routed against 27.3% control. Only 10% of routed threads went to the top tier. Their validation was the live A/B, and it worked because the users were their own engineers, who flagged a fast-only arm within a day. A customer-facing agent does not get that warning. So label before you route: pull 150 real threads, label each with the cheapest tier you would trust, write the criteria from the labels, and run the classifier against the set before live traffic. At Smartcat, 150 labeled translation requests gave engineering a working router in three days.

What LangChain shipped

The post is a clean piece of engineering writing, and it is specific.

They started from their own traces. A week of Open SWE threads, labeled by task type with an LLM classifier: new features were 22%, bug fixes 17%, test or no-op runs 16%. Cost and turn count per thread varied widely by type, which gave them the hypothesis that a lot of threads did not need the top model.

They picked three tiers along the cost and intelligence curve: fast, balanced, and performance. The router runs on the thread's first human message and has three parts, a base prompt telling the classifier to pick the least expensive model likely to complete the task, plain-language criteria per tier, and a classifier model.

They put it in the harness. Their reasoning is the line I would underline: the criteria are written for Open SWE's task set, the harness already has that context, and a generic gateway does not.

Then they measured. Half of threads routed, half always on the top model, 973 threads in total.

  • Merged PRs: 29.2% routed, 27.3% control, p = 0.49.
  • PR open rate: 38.9% routed, 39.6% control, p = 0.82.
  • Median cost per thread: $0.94 routed, $2.61 control. Mean down 42%, p90 down 37%.
  • Tier split: 56% balanced, 34% fast, 10% performance.
  • Median thread cost by tier: $0.097, $1.50, and $2.88.

They also ran the opposite control, router against fast-only, and ended it within a day. Engineers flagged the fast-only arm almost immediately because the output quality was disrupting their work.

The feedback channel you do not have

That last detail is the one I keep coming back to.

LangChain says offline evals are hard for a coding agent, because PR quality is hard to grade against a fixed dataset, and that an A/B test on live traffic works well when a dataset is too costly. For Open SWE that is a reasonable call. The users sit in the same Slack as the people running the experiment. The outcome, a merged PR, is counted on every thread. A bad arm got killed in a day by people typing "pretty expensive for this query" into a feedback box.

Now move the same design to an agent your customers use. The downside they name themselves, users subjected to a non-optimal router, lands on people who do not have your Slack. Very few of them rate a thread. The ones who got a worse answer from the cheap tier mostly do not tell you. You find out in the usage curve, and later in the renewal.

So the live A/B cannot be the first test. It has to be the second.

The drop: label before you route

I had a routing decision of my own at Smartcat, described in The Eval Is The Spec. The question was whether a translation request should stay with AI, go to a human linguist, or be flagged for escalation. Cost, quality, and turnaround all hung on it.

I spent an afternoon pulling 150 real requests from the last quarter and labeled each one with where it should have gone. I wrote the rubric: correctness of routing, cost efficiency, SLA compliance. Engineering took the set and had a working router in three days. We ran the eval daily, and every time the score regressed we could see which inputs failed and why.

Different decision, same shape. Here it is for model tiers.

  1. Pull 150 real threads from your traces. Sample across task types, the way LangChain's week of threads was split. Include the long ones.
  2. Label each with the cheapest tier you would trust with it. One column, three values. A second column for why, in a few words. The person labeling should be whoever reads the most output, usually the PM and one engineer.
  3. Write the tier criteria from the labels. The "why" column is your criteria, in your users' language. LangChain wrote theirs from the task analysis plus the model providers' guides.
  4. Run the classifier against the set. Before any live traffic. Count two errors separately: routed too low (quality risk) and routed too high (cost you did not need to spend). Fix the criteria until the first number is one you can defend.
  5. Keep the set. Rerun it on every model swap and every criteria edit. A new model on the fast tier is a one-line change in their middleware, and it is also a reason to rerun 150 rows.

Then run the A/B, on an outcome you can count per thread. Theirs was merged PRs. Yours is whatever your product is paid for.

This is the routing table from Gross Margin Is Your Job Now, with the eval threshold column filled in from data instead of from a meeting. About once a quarter I find a surface that is over-routed to the flagship for no reason other than that it shipped that way. LangChain just put a number on what that habit costs: on their task mix, nine routed threads in ten were sent below the top tier, and merged PRs held.

And it is the per-outcome line from Three Token Bills, One Unit. A 64% lower cost per thread is only good news next to the outcome rate for the same threads. They reported both. Report both.

The cluster this sits in, Enterprise AI Agents, makes one claim: what matters is what each successful outcome costs. A router moves the cost. The labeled set is how you know the outcome stayed put.

Related answer: How do you build a model router for an AI agent?

Sources: How to Build a Model Router in the Harness, Sydney Runkle and Eugene Yurtsev, LangChain blog, October 1, 2026. The Eval Is The Spec, falkster.com. Gross Margin Is Your Job Now, falkster.com.

Share this post

Also on Medium

Full archive →

Frequently asked

What did LangChain's model router save?+

Per Sydney Runkle and Eugene Yurtsev's 2026-10-01 post, routing Open SWE threads across three model tiers cut the median cost per thread from $2.61 to $0.94, 64% less, against a control that always used the top-tier model. The mean dropped 42% and the p90 dropped 37%. Across 973 threads, 29.2% of routed threads ended in a merged PR against 27.3% of control (p = 0.49).

Why does the model router belong in the harness and not a gateway?+

LangChain's argument is that choosing the right model takes domain and task context, the agent's prompt, tools, and task mix, which the harness already assembles and a generic gateway typically lacks. Their tier criteria are written for Open SWE's own task set, so the router is coupled to the agent on purpose.

How do you test a model router before it touches live traffic?+

Pull about 150 real threads from your traces, label each one with the cheapest tier you would trust with it, write the tier criteria from those labels, and run the classifier against the set. Keep the set and rerun it on every model swap and every criteria edit. At Smartcat, 150 labeled translation requests gave engineering a working AI-versus-human router in three days.

Is an A/B test on live traffic enough to validate a router?+

It was for LangChain, whose users are its own engineers and whose outcome, merged PRs, is counted per thread. Their fast-only arm was stopped within a day because engineers flagged the output quality. A customer-facing agent does not get that feedback channel, so the labeled set comes first and the A/B second.

How many requests actually need the most expensive model?+

In LangChain's test, 10% of routed threads went to the performance tier, 56% to balanced, and 34% to fast. Median thread cost was $0.097 on fast, $1.50 on balanced, and $2.88 on performance. Those are Open SWE's numbers on Open SWE's task mix, and your split comes from your own labels.

When does LangChain's router pick the model?+

Once, on the thread's first human message, and that model is used for the whole thread. They list mid-thread re-routing as future work and note the cost: switching models throws away the prompt cache.

THE SHORT ANSWER

PART OF

Enterprise AI Agents

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.