The day I read “Only 7% of turns needed a frontier model, and routing cut cost 74% for six points of accuracy,” I went looking for the source. The source turned out to be a Rust proxy sitting at NVIDIA-NeMo/Switchyard on GitHub, last pushed to today, with 2,475 stars and a v0.2.0 release that dropped August 10. The number wasn’t theoretical. It was measured against 145 multi-step agentic tasks in LangChain’s Deep Agents evaluation suite, with Opus 4.x running against a routed arm that mixed it with a smaller open model. This post is about what Switchyard actually does, where that 74% number comes from, and what it doesn’t fix.
The bet, in one sentence
A coding agent, a research agent, any multi-step LLM workflow, pays frontier prices for the capability ceiling of every turn, when most turns don’t actually need the ceiling. Switchyard’s bet is that if you give the agent a routing layer, pluggably teach it which signals mean “use the weak tier” and which mean “send to the strong tier,” and translate the wire formats on the way in and out, the frontier model becomes a fraction of requests, not a default. The cost saving is not magic. It’s the literal arithmetic of moving 93% of turns off the most expensive model in the deployment.
What Switchyard actually is
It’s a Rust binary, switchyard-server, plus a library crate, switchyard-libsy, that you can embed in your own gateway. The README is blunt about the maturity story: pre-alpha software, evolving rapidly, API and algorithms expected to change significantly before v1.0. The components are tagged individually: libsy is “Beta. Ready for trial integration.” switchyard-llm-client, switchyard-runner are “Alpha.” switchyard-server is a “Demo server, not for production use.” That last line matters if you’re thinking of wiring this into your team’s agent gateway tomorrow. It’s not ready for that yet. It is ready to read and learn from.
The interface contract is OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages — clients keep speaking their native API. Switchyard picks a configured backend, translates to that backend’s own format (Anthropic Messages in, vLLM native out, NIM out, Ollama out, anything OpenAI-compatible out), and translates the response back into the shape the client expects. That last part is non-trivial: streaming tool-call deltas, message-block reconstruction, system-prompt ordering, all need to round-trip cleanly across providers, and Switchyard ships a switchyard-translation crate dedicated to it. Without that, swapping a Codex-style client onto a self-hosted model is a Thursday you don’t get back.
Five routing strategies in one binary
The README’s strategy table is the right mental model:
- LLM Classifier: An extra cheap-model call decides whether this turn needs the weak or strong tier. The classifier itself becomes a request; you pay for it on every turn.
- Stage Router: No extra model call. Signals already in the conversation (tool results, error patterns, system-prompt shape) drive the decision. Where the classifier costs a round-trip, stage routing reads the trace it was given.
- Escalation Router: Weak tier runs first on every turn. A judge reads the weak-tier answer and escalates to strong only when the judge disagrees. Variant of
llm_classifierwithmode = "escalation". - Composite: An LLM classifier sets the tier a stage router falls open to. Nested routing. Useful when some signals you trust, some signals you want a model to read for you, and you want both in the same pipeline.
- Random: A fixed traffic split. For A/B tests, baselines, or cost experiments. Not a strategy you ship; a strategy you measure with.
The interesting move is the order: stage router and escalation router both avoid the always-on classifier cost. LLM Classifier is the one a lot of teams will reach for first because it’s the easiest to explain, and it’s also the one LangChain’s benchmark is implicitly measuring against something cheaper. If you imagine the routed arm in their 145-task run was a stage router rather than a classifier, the 74% number is closer to 80%, because you stop paying for the classifier call on the 93% that doesn’t need a frontier model. The benchmark post doesn’t say which router they used; the README gives you the search space.
Where the 74% number actually comes from
The LangChain Deep Agents evaluation suite is 145 multi-step agentic tasks. In the single-model baseline, every turn goes to Opus 4.x, paying Opus 4.x prices on every step. In the routed arm, most turns go to a smaller open model (a 30B-class model is the implicit reference, given the 8-point variance between it and a frontier model the LangChain team observed in the same suite), with Opus 4.x reserved for the 7% of turns that the router judges need it.
The arithmetic:
- 7% of calls on Opus carry 68% of the bill in the routed arm, because Opus is roughly ten times more expensive per token than the small open model, and the 7% of calls are also the longest calls (frontier turns are frontier turns for a reason).
- The remaining 32% of the bill is the small open model running on the other 93% of turns, plus router overhead. The open model gets no benefit from prompt caching, which LangChain notes as the second-largest line item — about a third of what Opus costs.
- Total bill: 26% of the single-model baseline. Six points of accuracy lost on tasks where neither model solved them well anyway.
That’s the 74% number: it’s not what routing adds, it’s what’s left when you stop paying for the 93% of turns that weren’t using the ceiling they were paying for.
Where this starts to bind to a real agent stack
Two things in Switchyard’s design map onto what’s already annoying about running an agent in 2026:
First, wire-format heterogeneity has been a silent cost for two years. Every agent harness — Claude Code, the Codex CLI family, Gemini CLI, Copilot, Cursor — picks one of three wire formats as the native one, and the moment you point it at a non-default backend you start hand-rolling translation. A proxy that does the translation once, at the protocol layer, means swapping the backend stops being a Wednesday the agent team owes you. Switchyard’s translation crate is where the actual engineering value sits; the routing algorithms are the reason NVIDIA is publishing this, but the translation layer is the reason a deployment team would use it.
Second, prompt caching breaks the naive router math. LangChain’s benchmark called this out: the small model doesn’t get cache reuse the way the frontier model does on repeat system prompts, and on long-running tasks that ends up dominating the routed arm’s bill. A real production router on top of Switchyard probably wants a tier in the middle — a model that can be served with a stable prompt prefix and isn’t the most expensive model in the deployment — to soak up the cache-ineligible traffic at lower per-token cost. The five strategies in the README are a starting point. The strategy that wins in production looks more like a three-tier router with cache-aware placement than a binary classifier.
That second point is the one Switchyard’s strategy list doesn’t (yet) have a primitive for. It’s a fair criticism: the five strategies are tier-selection primitives, not placement primitives. Placement — which instance of which model, on which accelerator, at which batching depth — is a separate problem and a separate router. Switchyard doesn’t try to solve it.
What the stage router is actually reading
I dug into docs/routing_algorithms/stage_router_routing.md via the repo tree to confirm. The stage router keys off three things it reads from the conversation state without making a model call:
- Tool result signals. Did the last tool call return an error? Did it return a value that the system prompt explicitly tells the model to handle internally? These are first-class routing signals; the small model can absorb them.
- System prompt shape. System prompts that look like “you are a planner, return a JSON” get treated as planning-stage turns; planners are usually small. System prompts that look like “you have full file access and need to write code that compiles” escalate.
- Conversation position. First turn and final-turn patterns. The first turn is often a clarifying question; the final turn is often a synthesis over many steps. Both can stay on the weak tier more often than you’d think.
This is the move that makes stage router cheaper than the LLM classifier: it’s pattern-matching against structured inputs, not asking a model to make a routing decision. The cost is that you have to write the patterns, and a pattern that’s too tight just behaves like the strong-tier-everywhere default. The cost is that you have to maintain them as the harness changes. The win is that there’s no model in the loop paying per turn to decide routing.
Trade-offs and what Switchyard doesn’t fix
Switchyard doesn’t solve the small-model capability ceiling. If the 30B-class model genuinely can’t solve 12% of tasks the frontier model can, the routed arm gives up those tasks. The 7% frontier-call number already reflects that — the router isn’t claiming every task is solvable on the weak tier. If your task distribution is bimodal (“questions about X” and “questions about Y” with no overlap in solver), a stage router that doesn’t know about that distribution will route both the same way, and you’ll either over-spend on X or under-spend on Y. Knowing your distribution matters.
Translation fidelity is not free. The wire-format layer has bug surface. LangChain’s translation layer, LiteLLM’s adapter layer, every adapter in this space has lost time to streaming delta ordering and tool-call block reconstruction. Switchyard’s README is honest that switchyard-runner is alpha; expect translation bugs against new model releases for at least another quarter. Running this in production means budget for a person to debug the proxy when a new model drops a slightly different message block shape.
The repo doesn’t try to solve batching and placement. vLLM gives you PagedAttention and continuous batching. TGI gives you a different batch scheduler. SGLang gives you radix attention. None of that crosses through Switchyard — Switchyard hands a request to a single configured upstream URL and waits for the response. If your frontier tier needs to be on a shared inference cluster with twenty other workloads, Switchyard doesn’t know. A real production deployment has the routing layer plus a placement layer; Switchyard is the first.
It’s pre-alpha. The README’s Maturity section is not a disclaimer for lawyers — it’s the operational state of the project. v0.2.0 has been out for sixteen days. The API will change. If you’re embedding switchyard-libsy into an agent runtime today, plan to rewrite against v0.4 or v0.5 before v1.0. Read the source, fork it, stay close to the diff. Don’t pin it.
The 7% / 74% / 6-points result is one benchmark. LangChain’s Deep Agents suite is a specific distribution of multi-step coding and research tasks. It is a real distribution, not a contrived one, but it is not your distribution. Run your own before you commit budget to a router.
Where I’m landing
The reason this matters even before Switchyard hits v1.0 is what it says about the cost shape of an agent. The naive shape is “every turn to the most capable available model.” The measured shape in 2026 is “most turns don’t need the most capable model, and the cost of finding that out is not paying for it.” That gap isn’t unique to NVIDIA. LiteLLM has been doing cheaper variants of this for a year. Portkey and OpenRouter are doing it for paid customers. What Switchyard contributes is a published open-source reference implementation in Rust, with a configured strategy vocabulary that lets you pick which router primitive you want instead of inheriting someone else’s choice. That’s worth a fork, even at pre-alpha.
The question I haven’t answered is which of the five strategies wins on my own workload. I don’t know yet. The right next thing to do is put the stage router against an eval suite I trust and measure 7% / 74% / 6-points the same way LangChain did. If the number holds on my distribution, the next move is the three-tier extension — small / cache-stable-mid / frontier — and a translation-fidelity test suite that catches the proxy when a new model release shifts a message block.
That’s a different post. For today, the headline is that the routing primitive moved from a vendor implementation to an open-source one, and the open version names what the vendor versions usually don’t: the bill is going to look like this and only this on this workload.
References
- NVIDIA NeMo Switchyard on GitHub — Rust proxy and library, 2,475 stars as of today, v0.2.0 tagged August 10
- NVIDIA blog post launching Switchyard alongside Nemotron 3.5 Lightning — August 11
- LangChain benchmark: 145 Deep Agents tasks, 7% / 74% / 6 points — Srimanth Tangedipalli and Karan Singh, August 11
- NVIDIA Developer blog: routing with Switchyard — covers the five routing stages
switchyard-libsycrate docs — embed-the-router path for non-server deployments- vLLM PagedAttention — Kwon et al., SOSP 2023 — the inference engine that’s the natural upstream for Switchyard’s router
- LiteLLM — the Python adapter whose translation layer Switchyard’s design is the Rust counterpart of
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.