How do you build a model router for an AI agent?

THE SHORT ANSWER

Put the router in the agent's harness, where the task context already lives, and build a labeled set before it sees live traffic. LangChain's router for Open SWE (2026-10-01) picks one of three model tiers on a thread's first message and cut median cost per thread from $2.61 to $0.94, 64% less, with merged PRs flat across 973 threads. They validated it with a live A/B on their own engineers. For a customer-facing agent, pull 150 real threads, label each with the cheapest tier you would trust, write the tier criteria from the labels, run the classifier against the set, and rerun it on every model swap. At Smartcat, 150 labeled translation requests gave engineering a working router in three days.

LangChain's post is the build guide: understand the tasks from traces, pick models along the cost and intelligence curve, put the router in the harness, and track task outcomes. Label Before You Route adds the step that makes it safe when the users are customers and not your own engineers, a labeled set of 150 real threads that the classifier has to pass before live traffic.

The method is the one in The Eval Is The Spec, where 150 labeled translation requests at Smartcat replaced a routing PRD, and the artifact it feeds is the routing table in Gross Margin Is Your Job Now: model, fallback, and eval threshold per surface, reviewed quarterly. Both belong to the Enterprise AI Agents argument about what each successful outcome costs.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-02 · 1 min read