$ cat wiki/concepts/model-routing.md
Model Routing
Definition
Choosing which model answers which request — or which step of a request — at runtime, rather than binding an application to one model. The decision is made by a component sitting between the caller and the models: a classifier, a heuristic, or a small trained router.
Two distinct motivations produce the same machinery, and this page keeps them apart because they fail differently:
- Cost routing — send the cheap step to the cheap model, escalate only what needs it. Failure mode: a task is under-served and quietly gets a worse answer.
- Safety routing — send a request that trips a classifier to a more restricted model. Failure mode: a legitimate request is downgraded. Anthropic calls this a fallback (source).
Why It Matters
Routing is where the published price of a model stops describing what a workflow costs. Once a router is in the path, the per-token rate on a model page is an input to a cost function, not the cost — and the wiki's own spec tables cannot express that.
It also relocates a safety decision. When a classifier picks the model, the classifier, not the model, is the thing that determines what a user is allowed to get — and it is usually the component with no model card, no benchmark and no version history.
State of the Art (as of 2026-08-24)
Routing measured from the demand side, and it caps what a frontier model can charge (2026-08-23)
Everything else on this page is supply-side — routers, gateways, classifiers, the infrastructure that makes routing possible. The Ramp AI Index for August 2026 is the first thing this wiki holds that measures whether businesses actually do it, and what it costs the model at the top.
Ramp computes the index from corporate card and bill-pay spending across roughly 70,000 businesses. Two months after launch, Claude Fable 5 accounts for 11.4% of Anthropic dollar spend and 6% of tokens, and Claude Opus 5 — launched in late July at $5/$25 per million tokens, half Fable 5's $10/$50 — has already overtaken it in enterprise spending (Ramp, speaking to the Financial Times). Fable 5 generated about 75% as much July revenue as OpenAI's flagship at roughly double GPT-5.6 Sol's per-token price (source).
The behaviour Ramp describes is routing stated as procurement policy: cheap models for routine company-wide work, the top tier reserved for the small number of cases where failure is expensive. Ramp explicitly frames this as a change — earlier, enterprises were more likely to migrate onto the most capable model available. Its own summary: "With Fable 5, we've found a new upper bound for how much businesses are willing to spend on AI."
Why this belongs on this page and not only on a model page. The wiki's July
digest recorded that price stopped being a tier and became the argument. This is the
mechanism underneath that: once routing is in the path, a frontier model is not
competing for a workload, it is competing for the fraction of a workload a router
escalates to it — and that fraction is small, set by someone else's cost policy,
and invisible in any benchmark. A model page's Pricing row describes a rate; it
cannot describe a ceiling on volume that the rate itself creates.
The second named drag is a safety policy. Ramp economist Ara Kharazian attributes the disappointment to "price + data retention requirements" — the 30-day retention Anthropic requires on Mythos-class traffic. That is the first measurement this wiki holds of Safety Monitoring and Data Retention showing up as a commercial cost, and it is recorded there rather than repeated here.
What this measurement does not establish. It is one vendor's customer base, not the market; "Anthropic dollar spend" is not broken out into API versus subscriptions; and nothing read controls for the fact that Fable 5 was export-suspended from 2026-06-12 and returned credits-only. A model unavailable for part of its own launch window is not a clean adoption measurement, and the absence of that control is a real limit on the conclusion, not a quibble.
The best-known router is acquired by a payments company for $7B+ (2026-08-17)
Stripe acquired OpenRouter for more than $7 billion, against a $1.3 billion valuation set by a $113 million Series B three months earlier. Reporting puts earlier talks near $10 billion and attributes the ~30% fall to summer-2026 model price declines — a single outlet's attribution, recorded as reported (source).
OpenRouter is the most widely used instance of the pattern this page describes: a single endpoint in front of many providers, choosing per request by capability, price and availability. What the price says about the pattern is the part worth recording — a 5× step-up in three months values the routing layer, not any model behind it, and the buyer is a payments company. The two structural readings both fit and nothing read separates them: routing is being valued as metering infrastructure (who called which model, at what cost, billed to whom), or as distribution (the gateway chooses, so the gateway has the leverage). Neither is stated by either party.
This wiki is not a neutral observer of this transaction. scripts/spec-check.py
reads OpenRouter's public catalogue as the daily cross-check on every Pricing
and Context window cell in this wiki's model pages, and CLAUDE.md names the
catalogue as the source rather than a cross-check for models whose vendor
publishes no list price. An ownership change at the catalogue is a change to this
repository's verification chain. Nothing read states any intended change to
the catalogue, its public API or its pricing data — recorded so that if one comes,
the date it was foreseeable is on the record
(source).
NVIDIA NeMo Switchyard — routing shipped as open-source infrastructure (2026-08-11)
NVIDIA released NeMo Switchyard, an open-source model routing library for AI agents, on the same day as Nemotron 3.5 Lightning. It routes each prompt to a model per step of an agent workflow (source).
It ships two families (source):
- Tuning-free routers, which decide without training on workload-specific data: an LLM classifier, a stage router, and an escalation router
- Tunable routers, trained on the workload
NVIDIA's internal benchmark claim: frontier-level accuracy at nearly one-third the task-completion cost of Opus 4.8 alone (source) (VentureBeat).
Read the pairing, not just the library. NVIDIA shipped a cheap model and, the same day, the software that decides when to use a cheap model — with the comparison baseline being a competitor's flagship. The router is the argument for the model.
Anthropic's biology fallback — routing as a safety control (2026-08-07)
Claude Fable 5 routes a request to a less capable model — Claude Opus 5 — when a classifier judges it to touch safeguarded biology. Anthropic retrained that classifier's constitution and reported ~85% fewer biology-related fallbacks in testing, with the reduction in total fallback volume spread very unevenly across surfaces: ~67% Claude.ai, ~55% Cowork, ~17% Claude Code, ~7% Claude Platform (source).
That spread is the most informative routing measurement this wiki holds. It says the router was firing on consumer questions far more than on developer traffic — i.e. the misrouting was concentrated in one population, invisible in the aggregate.
Latency-and-capacity routing in agent platforms (2026-06)
Azure Agent Mesh, announced at Microsoft Build 2026, performs automatic routing based on latency and GPU availability across on-prem Windows, Windows 365 and Azure Arc edge, with a GA target of Q4 2026 (source). This is the third motivation — routing for placement rather than for cost or safety — and it predates the other two on this page.
2026-08-19 — an audit of what a hosted route actually serves you
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391) formalises hosted model routing as a stochastic process and audits it black-box — no probability information required from the target (source):
| Statistic | How it is obtained |
|---|---|
| AFL — average fidelity loss | a frozen constrained context sent repeatedly; an output distribution reconstructed from returned text counts; a null-bias-corrected within-window mean coarsened-KL |
| EFL — extreme fidelity loss | independent runs; the empirical upper tail of a run-level reference-centered-surprisal statistic |
| Reported: AFL shows strong linear descriptive agreement with a logprob-derived | |
| comparator across three logprob-capable route conditions; 20-run probes across | |
| seven route snapshots reveal route-specific EFL variation; and the split | |
| that matters — **neither statistic has much detectable route-level association with | |
| GPQA-Diamond accuracy**, while **pronounced EFL coincides with a declining | |
| Terminal-Bench pass rate as task exposure increases**. |
This is the first instrument on this page that measures the routing layer
rather than the routing decision. Everything above is about choosing a model;
this is about whether the model you chose is the one answering. It arrives two days
after Stripe acquired OpenRouter for more than $7 billion, and it names the
exposure that acquisition made visible: a routed request's provider is not the
vendor, which is the same finding that forced this repo's own
scripts/spec-check.py to compare against first-party endpoints rather than
headline catalogue prices.
The GPQA result should change how a third-party figure is read here. A route can show pronounced fidelity loss and score normally on a single-turn knowledge benchmark — and single-turn-shaped numbers are what most third-party figures in this wiki are. It bears directly on Qwen 3.8 27B's Artificial Analysis Intelligence Index of 52, recorded today: a leaderboard figure is produced through some serving path, and nothing states which.
What it does not give: no provider or route is named, no AFL or EFL value is published, and the Terminal-Bench decline has no magnitude. The method also detects divergence without identifying its cause — quantisation, batching, speculative decoding, sampling drift and outright substitution all fit, and nothing read separates an engineering trade-off from a swap.
2026-08-29 — the switch itself has a price, and the right interface reverses with direction
Every claim on this page so far treats the act of routing as free: the cost model is the price difference between two models, and the quality model is the capability difference between them. The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents (arXiv:2608.24358) is the first measurement in this wiki of the switch itself.
The setting is a long coding-agent run — "dozens of model calls, tool uses, and code edits" — and the decision is the practical one: escalate to a stronger model when a cheap one struggles, or downshift once the hard reasoning is done. Pairs of low-cost/low-capability (LC) and high-cost/high-capability (HC) models from the Claude and GPT families were run across three variables — direction, timing and interface — with full-trajectory transfer, compaction and trajectory removal compared while preserving the repository state (source).
Holding the repository constant is what makes it a clean experiment: the only thing varied is what the receiving model is told about how it got there.
- Escalation is the expensive direction. Full-trajectory LC → HC "recovers less than half of the LC-to-HC quality gap while incurring a substantial cost premium". The paper names this the handoff tax.
- Downshift "offers a favorable cost-quality point."
- The preferred interface reverses. "Reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality."
The asymmetry is the finding. A weak model's reasoning trace is actively misleading to a stronger model, which does better with less of it; a strong model's trace is load-bearing scaffolding for a weaker one, which does worse without it. The same artefact is a liability in one direction and an asset in the other, so no single context-handoff policy can be correct for both, and a router that carries "the conversation so far" uniformly is wrong half the time by construction.
Two limits. The paper publishes no absolute numbers — "less than half" and "substantial cost premium" are the stated forms — so the tax cannot be weighed against the price gap that motivates the switch, which is the decision it is about. And compaction, the practically interesting middle option between carrying everything and carrying nothing, has no separately reported result in anything read.
2026-09-11 — the router is sold as a model, with a per-token price and a benchmark table
Everything above treats routing as infrastructure: a library, a gateway, a classifier, a $7B acquisition of a catalogue. Sakana AI sells it as a model. Fugu Max and Fugu Ultra v2, both released 2026-09-11, are two tiers of one learned orchestrator behind a single OpenAI-compatible endpoint, priced per token and benchmarked against models. Sakana's own phrase is "a Multi-Agent System, Delivered as One Model", reaching its results "by dynamically coordinating and orchestrating a diverse pool of powerful models" (source).
| Fugu Max | Fugu Ultra v2 | |
|---|---|---|
| Pricing | $2/M input · $6/M output | $5/M input · $30/M output · $0.50/M cached |
| Context window | 1,000,000 | 1,000,000 |
| Claimed result | best overall on six benchmarks; Pareto frontier expanded on 7 of 10 | best or joint-best on five of eight; Chartography 48.3, DeepSWE 74.3 |
| Pool | widened to open-weight and specialised models, including NVIDIA Nemotron | excludes Fable 5, Fable 5.1 and GPT-6 Astra |
| This page's third Open Problem stops being hypothetical. It reads: *"If | ||
| frontier-level accuracy is achieved by a router over several models, the number | ||
| characterises the system."* That was written about NVIDIA's internal benchmark | ||
| claim for a library. It now describes a **published spec table on a model page in | ||
this wiki** — two of them — where Pricing, Context window and a benchmark | ||
| column all describe an orchestrator over a pool that **is never enumerated in | ||
| anything read**. |
The pool claim is the new thing, and it is sharper than a benchmark win. Sakana states that Ultra v2 reaches its scores "without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool" and "does not rely on individual proprietary frontier models to deliver frontier output". A router beating a frontier model is unremarkable — it can call that model. A router beating three frontier models it has removed from its own pool asserts that the capability lives in the coordination. No ablation against a v2 run with the three restored appears in anything read, and no orchestrated baseline over the same remaining pool is reported, so what the conductor contributes over a simpler assignment policy is unmeasured.
The first Open Problem below is also untouched by it. Neither release publishes a misroute rate. Sakana publishes what the system scored, which is an outcome, and the same limitation that survived xRouteBench survives this.
Two things this page's own machinery should note. Fugu's per-token price is
charged on the caller's tokens while the answer is assembled from other vendors'
models at their rates — so this is the first Pricing cell in this wiki that
cannot describe what a request costs to serve, and nothing read says who absorbs
the difference. And SWEFish, one of the six benchmarks Fugu Max leads, is
Sakana's own internal benchmark, which is the
Eval Harness Configuration problem arriving inside a routing claim.
2026-09-24 — the routing decision moves inside the system, and the router's signal is its own uncertainty
Two papers in one HuggingFace Daily Papers snapshot take the argument below — that the decision layer should not be a language model — and apply it inside systems rather than at their edge (source).
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure — a cascade routed by confidence. A decision-only judge is compared against sixteen generative and reward-model judges under blinded human adjudication, landing within three percentage points of the strongest comparator on ordinary preference and evidence-grounded factuality at 0.36% of that comparator's fee. Its gap to that comparator is concentrated in low-confidence decisions, so a frozen cascade — accept confident verdicts, escalate uncertain ones — retains 99% of the comparator's accuracy at lower cost.
This is a routing architecture every entry above lacks. The routers on this page decide before the work, from the request: a price cap, a learned classifier, a per-token product. Here the cheap model does the work first and routes on its own uncertainty about the answer it just produced. The routing signal is generated by the thing being routed, which removes the standing problem that a router must predict difficulty it cannot yet observe.
It also relocates where the claim can fail. A confidence-routed cascade is only as good as its calibration, and calibration is the one property Jev has no published number for — the page records RLCD as named and unspecified. The 99% figure is evidence the confidence signal carries information; it is not a calibration measurement.
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents — five routing decisions inside one agent's memory. A System-One control plane takes over query routing, retrieval-budget allocation, graph traversal, candidate scoring and adaptive stopping, leaving System Two for reasoning and answer synthesis. LoCoMo LLM-as-a-Judge 0.777 (+11.0% relative), memory construction 158 s (6.6× faster), average query latency 0.93 s (−36.7%).
Here the router is a component, not a product. Every prior entry on this page routes between vendors' models across an API boundary; this routes between planes of one system, and the thing being economised is generation on a critical path rather than a per-token bill.
What neither paper establishes is who wrote it. Neither snapshot entry
carries an author list or affiliation, arxiv.org answers EGRESS_BLOCKED from
this run's sandbox, and whether TypeSafe AI authored or funded
either is unknown and is not asserted. The strongest comparator is not
named in either, and no benchmark is named in the judge paper.
One figure does move against the vendor. TypeSafe's own cost claim is
~4,000× less per call; 0.36% of a fee is ~278×. The two measure
different units and neither is adopted, but the 14× spread is recorded on
Jev under ## Conflicting Reports.
2026-09-15 — the argument that the router should not be a language model at all
Every entry above takes for granted that the thing deciding where a request goes is an LLM, or is measured against one. TypeSafe AI left stealth on 2026-09-15 arguing that it should not be. Its first model, Jev, is pitched at a class it calls System One Models, whose stated jobs are exactly four — decide, classify, route, score — and which do not generate text at all: structured or natural-language state in, a typed value out (a choice, a score, a calibrated probability), produced in one parallel forward pass rather than token-by-token (3 passes) (source).
| Figure | Value | Passes |
|---|---|---|
| Input price | $0.042 per million tokens | 3 |
| Output price | free (stated: too inexpensive to meter) | 3 |
| End-to-end latency | 70–500 ms | 3 |
| Training method | RLCD — Reinforcement Learning for Calibrated Decisions | 2 |
| Accuracy vs "the most expensive frontier models" | within 3 points, on TypeSafe's own workflow evaluations | 1 |
| Cost per call vs the same | ~4,000× less | 1 |
| Funding | $40M seed, led by DCVC | 2 |
| Why it lands on this page rather than only on a model page. The 2026-08-23 | ||
| entry above established, from the demand side, that **price and not intelligence | ||
| decides most routed traffic**, and that this caps what a frontier model can charge. | ||
| The 2026-09-11 entry recorded the first router sold as a model with a per-token | ||
| price. This is the next step in the same argument and a more aggressive one: if the | ||
| decision is worth $0.042 per million tokens, then routing was never a frontier | ||
| workload and the price cap is not a cap but a floor **two orders of magnitude | ||
| lower** than anything this page has priced. For scale, Fugu Max — the | ||
| router-as-a-model — charges $2/M input, about 48× more. |
The claim is unverified in every particular that matters. There are no
independent evaluations at launch (2 passes). The comparator is "small frontier
LLMs" and is never named, so every multiplier is a ratio with one side
missing. RLCD is named but not specified — reward function, architecture,
training procedure and calibration methodology all undisclosed — and
calibration, the single property it is stated to optimise, has no published
number at all. And TypeSafe publishes four different answers for its own speed
and cost multipliers across its own surfaces (193.6×/444.6× on the home page,
40–200× in the blog, 20–200×/40–400× in the founder's thread, >100×/>200× in the
AINews headline). That is not the provider-spread problem CLAUDE.md documents for
spec-check: there is one provider here, publishing four figures. The table is
on TypeSafe AI and no figure is adopted.
Open Problems
- Nobody publishes the router's error rate. Partly addressed, 2026-08-19 — Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391) measures a different error than the one named here: not "was the wrong model chosen" but "did the chosen model's output distribution shift". Both are unmeasured elsewhere; neither substitutes for the other. Every claim on this page is about the outcome of routing (cost saved, fallbacks reduced). None states how often the router sent a request to the wrong model. Anthropic's per-surface spread is the closest thing to that number and it is an inference, not a measurement. Unchanged by the field's first benchmark: xRouteBench scores quality and cost, which are also outcomes (source).
- No benchmark covers safety routing. The five xRouteBench task families (generic, memory-augmented, vision, time-series, personalized) are all capability-and-cost. The fallback case Anthropic ships optimises something else entirely, and appears in none of them — so the half of this page with a live safety consequence still has no measurement at all.
- A routed benchmark result is not a model result. If frontier-level accuracy is achieved by a router over several models, the number characterises the system. This is the same reporting problem Eval Harness Configuration records for scaffolds, arriving now for model selection. No longer hypothetical, 2026-09-11 — Fugu Max and Fugu Ultra v2 publish exactly such numbers on model pages, against an unenumerated pool, and one of the six benchmarks Fugu Max leads is the vendor's own.
- Cost claims are baseline-shaped. "One-third of Opus 4.8 alone" compares against running the most expensive available model for every step — a baseline nobody chooses. What it saves against a competently hand-tuned assignment is unstated.
- Safety routing has no appeal path. A user downgraded by a classifier is not told which component made the decision, and the classifier has no published version.
- Router-level attacks are unexamined here. If routing decisions are driven by a classifier reading the prompt, the prompt can be written to steer it. No source read addresses this.
Key Papers
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867) (2026-08-16) — the first research this page has read, and it goes after the thing the Open Problems above say is missing: comparability. It recasts routing as a sequential decision process with five named components (context encoders, model encoders, scoring functions, decision rules, learning signals), and ships xRouteBench, which scores routers jointly on response quality and inference cost, plus an open-source library of 16+ routers. Learned routers beat the strongest fixed-model baseline by 14.6% relatively (source).
Two things it does and does not settle. Evaluating quality and cost together is what makes an under-serving router legible at all — on cost alone, sending everything to the cheapest model wins. But no error rate appears in it, so the first Open Problem below survives its own first paper: xRouteBench measures outcomes, not misroutes. And nothing read names which 16 routers are included, so it cannot yet be pointed at NeMo Switchyard.
Stealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867) is adjacent but not about routing: it turns on reasoning traces being interchangeable across models within one provider, which is the same substitutability that makes routing possible.
Referenced by
Sources
- sources/papers-daily/hf-daily-2026-09-24.md
- sources/blogs/typesafe-2026-09-15-jev-system-one.md
- sources/blogs/sakana-2026-09-11-fugu-max-ultra-v2.md
- sources/papers-daily/hf-daily-2026-08-28.md
- sources/blogs/ramp-2026-08-23-ai-index-august-2026.md
- sources/papers-daily/hf-daily-2026-08-19.md
- sources/blogs/stripe-2026-08-17-openrouter-acquisition.md
- sources/blogs/nvidia-2026-08-11-nemotron-3-5-lightning-switchyard.md
- sources/blogs/anthropic-2026-08-07-fable5-biology-safeguards.md
- sources/blogs/microsoft-2026-06-01-build-project-polaris.md
- sources/papers-daily/hf-daily-2026-08-16.md