The interesting thing about Sakana Fugu Ultra v2 is what it does not use. The Sep 11, 2026 release from Sakana AI is a learned multi-agent orchestrator trained to assemble, route, and coordinate a swappable pool of expert agents — and its production agent pool, the one that hit Chartography 48.3, DeepSWE 74.3, and SWEFish (best overall), explicitly excludes Fable 5, Fable 5.1, and GPT-6-Astra. The flagship frontier models. The ones every other orchestrator routes to. Fugu Ultra v2 has them turned off in its own runtime.
That single constraint is what makes this release a real architecture story and not another “we wrapped the API in a router” post. The Sakana release page says it plainly: “Fugu Ultra v2 achieves these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool. Fugu Ultra v2 does not rely on individual proprietary frontier models to deliver frontier output.” That’s a bet — and it’s the kind of bet that either falls apart on the benchmarks or shifts how you think about what an “orchestrator” is.
What a learned orchestrator actually is
Most multi-agent systems are hand-designed graphs. You decide that step 2 is a critic, step 3 is a tool-caller, step 4 is a verifier; you write the conditional edges; you commit it to a LangGraph or a CrewAI or a fixed pipeline. The release page for TradingAgents v0.4.0 (which I wrote about this morning) is a perfect example — four analysts feed a Bull/Bear debate, the Trader proposes an action, three risk debaters grade it, the Portfolio Manager decides. Every node is named, every edge is fixed. The graph is the product.
Fugu is the opposite. The same release page describes the architecture: “Fugu models are themselves language models trained to understand user queries and dynamically devise agentic scaffolds to solve them.” The orchestrator is a language model. You give it a query, and it has to plan which sub-agents to call, in what order, with what prompts, and how to verify the outputs. The model writes the workflow at inference time, per query, as a sequence of natural-language coordination steps.
This is grounded in two ICLR 2026 papers Sakana published earlier in the year: TRINITY and the Conductor. TRINITY uses a lightweight evolved coordinator — evolutionary algorithms, not gradient descent — to assign Thinker / Worker / Verifier roles across turns. The Conductor uses reinforcement learning to discover natural-language coordination strategies, the actual prompts and communication patterns, that help diverse LLM pools outperform individual workers on hard reasoning benchmarks. The release notes that the v2 models combine large-scale fine-tuning, evolutionary algorithms, and RL — so both papers are in the production stack, not just one of them.
The bet, in one sentence
You don’t need the closed frontier models to deliver frontier output. You need a learned policy that picks the right mix from a swappable pool of open and specialized models — and you need to train that policy on real benchmark signal, not on synthetic “good workflow” demonstrations.
The pool is swappable in two senses. First, end users can opt specific agents out — the release page explicitly markets compliance/privacy/sovereignty use cases. If a regulated industry can’t have certain providers in the pool, you can drop them and Fugu keeps working with what remains. Second, Sakana has the option to swap a model in their own pool — replace Qwen 3.x with DeepSeek V4.1 Flash, replace Nemotron with whatever comes next — without retraining the orchestrator. The orchestrator’s job is to compose. If the components move, the orchestrator’s policy should still generalize.
That decoupling is the architectural payoff. A monolithic frontier model is a single point of failure: if your provider revokes your API access, raises prices, or gets geopolitically cut off, your agent stack breaks. A learned orchestrator with a swappable pool absorbs any single change. Sakana is explicit about this in the release: “orchestration is not a trade-off between cost and capability. It is the architecture that optimizes both simultaneously.” The pitch is resilience, not just performance.
What’s actually hard on the benchmarks
The Sep 11 release lists results on eight benchmarks. Three numbers make the architecture argument land:
- Chartography 48.3 vs Opus 5 27.3 vs Fable 5 29.5. Visual reasoning and structured-data interpretation. Fugu Ultra v2 wins by 21 points over Opus 5.
- DeepSWE 74.3 vs models that cost 3–5x more per token. Real-world software engineering, the kind of multi-file refactor + verification task that breaks a one-shot LLM.
- Toolathon, SWEFish, GDP.pdf, Terminal Bench 2.1, GPQAD, AA-LCR, AutomationBench — best or joint-best on five of eight, top-2 on seven of eight.
The Chartography gap is the most interesting. Opus 5 is the closed frontier multimodal; the gap is not a 1.5x improvement, it’s a 1.77x improvement. If Fugu Ultra v2 were routing to Opus 5 internally, you’d see something close to the Opus 5 score, maybe a few points above. The 21-point gap means the orchestrator is finding combinations of specialist models (vision-language, structured-data, code-verifier) that beat what any single model in the pool does alone. That’s the learned orchestrator doing what its training taught it to do.
The cost angle on Fugu Max is sharper. $2 input / $6 output per 1M tokens is 40–60% lower than Sonnet 5, GPT-5.6 Terra, and Kimi K3. The release says Fugu Max sits on the Pareto frontier of cost-performance across seven of ten benchmarks — meaning it gets you “near-frontier performance at a fraction of frontier cost” without sacrificing the agentic capabilities. If you’re running a 24/7 production agent and the frontier model is your bottleneck, Fugu Max is the obvious migration target.
The mechanism, in one paragraph
The orchestrator reads the query, samples a coordination plan as a sequence of natural-language steps (which agents to call, what to ask each, how to verify), executes the plan against the swappable pool, observes the outputs, and writes a final response. Each step in the plan is a token sequence — Fugu is an LM, so the plan lives in the same latent space as any other generated text. The training pipeline combines large-scale fine-tuning on synthetic orchestration traces with evolutionary search over candidate coordinator policies (TRINITY-style) and RL on hard-reasoning benchmarks where the reward signal is the final answer quality (Conductor-style).
What makes this work in production — and what I think most “we built an orchestrator” posts skip — is the swappable-pool interface. Fugu doesn’t bind to specific model checkpoints; it binds to a registry of available agents with capability tags. When you opt a model out of the pool, the orchestrator’s policy has to discover, at inference time, which combinations still work. That’s the load-bearing design choice. A monolithic router trained on (query → GPT-6) would collapse the moment GPT-6 is removed; a learned orchestrator trained on the abstract structure of agent capabilities re-plans. Whether it actually does that well is the open question the benchmarks only partially answer.
What the numbers actually show, and what they don’t
Fugu Ultra v2 wins on the hard benchmarks where frontier models already win. That’s a useful signal — it means the orchestrator is preserving the capabilities of the underlying models, not degrading them. Where it loses, it loses on benchmarks that reward raw monolithic reasoning: pure closed-book knowledge, simple arithmetic, conversational coherence — anything where one good LM beats three OK ones cooperating.
The release is also honest about what Fugu Ultra v2 is not. It doesn’t expose the agent pool publicly — you can’t audit which exact models are in the pool, what their system prompts are, or how the orchestrator decides to call them. Sakana calls this “Fugu is an architecture, not a single model,” which is true and also means you have to take the orchestrator on trust. The reverse-engineering community (Requesty published a June 2026 deep-dive on Fugu Ultra) suggests the conductor layer does query classification first, then decides how many models to call. The v2 release doesn’t materially change that surface, just the underlying policy.
Fugu Max has a different tradeoff. It expands the agent pool to its largest size and dynamically routes to the leanest model capable of solving each task — which means the cost saving depends on how often the orchestrator picks a small model. If your workload is dominated by queries that need the full pool, the savings evaporate. The Pareto-frontier claim is true on average across the eight benchmarks; it’s not a guarantee for any individual query.
Trade-offs and what it doesn’t fix
Three honest limits.
Closed-source orchestrator on top of open models. Elie Bakouch’s tweet (June 2026) hit this cleanly: “to be clear, this is a closed source orchestrator on top of [open models].” The swappable-pool pitch sounds like open-weights sovereignty, but the orchestrator itself is a closed Sakana model. You can audit the agents you put in the pool; you can’t audit Fugu’s policy that decides which ones to call. If your threat model includes “the orchestrator is silently biased toward Sakana-preferred agents,” you have to take that on faith. The technical report (arXiv:2606.21228) describes the training paradigm in detail but doesn’t ship the orchestrator weights.
Latency cost compounds with pool size. Fugu Ultra v2 has to do more forward passes than a monolithic model. For a query that needs three sub-agents with verification rounds, you’re looking at 6–10 LM calls per answer. On Fugu Max with a large pool, the orchestrator’s own classification pass adds latency before any sub-agent runs. For real-time chat the base Fugu model (the third tier, not in this release) is the latency-balanced option; Fugu Ultra v2 is the slow option. The release is explicit: “Fugu Ultra v2 trades additional latency for higher quality.” That’s the deal.
The pool composition is the real product. Fugu Ultra v2’s scores depend on having the right agents available. Sakana’s pool includes NVIDIA Nemotron family models (the Sep 2026 NVIDIA partnership is integrated) plus whatever open-weights and specialized models they keep current. If you’re a customer running Fugu and a key agent in the pool gets deprecated, your output quality changes — not because the orchestrator got worse, but because the pool did. The architectural promise of “swappable” assumes Sakana is keeping the pool well-curated, which is a long-term operational bet you can’t verify from the release notes.
What to actually do with this
The interesting move isn’t “switch your production agent to Fugu Ultra v2.” It’s recognizing what the release is claiming: a learned policy that can compose open models can match or beat the closed frontier. If Sakana can do this in September 2026 with TRINITY + Conductor, the bet is that any frontier lab can do this — and any team that’s been hand-designing multi-agent graphs (TradingAgents, ACO System, my own orchestration code) is competing against a learned policy that will only get better. The 87.6% SWE-bench Pro number from earlier this year was a monolithic capability story. The Fugu Ultra v2 Chartography number is a composition story.
The open question I’d want answered before committing: does the orchestrator’s policy transfer across pools? If I take Fugu Ultra v2’s API endpoint and replace every agent in the pool with a different open-weights model — say, swap Nemotron for Qwen3-8 27B, swap the vision model for InternVL 3.5 — does the orchestrator still hit Chartography 48.3, or does it fall back to the average capability of the pool? The release doesn’t show that experiment. If the orchestrator generalizes, the architecture is real. If it’s overfit to the curated pool, Fugu is just another closed-source wrapper. The benchmarks alone don’t tell you which.
The Sep 11 release leaves the open question open. That’s the honest place to leave a post about a learned orchestrator.
References and where to dig further
- Sakana Fugu Max / Fugu Ultra v2 release page — the headline numbers, the pool-exclusion note, the Pareto-frontier framing
- Sakana Fugu product page — the four-tier model lineup (Fugu, Fugu Ultra, Fugu Max, Fugu Cyber), the TRINITY + Conductor grounding
- Sakana Fugu Technical Report (arXiv:2606.21228) — the full training paradigm (fine-tuning + evolutionary algorithms + RL)
- Requesty: Inside Sakana Fugu Ultra (June 2026) — the reverse-engineered conductor/orchestrator architecture
- Decoding Sakana Fugu Technical Report (Outcome School) — a readable walk-through of the same report
- TRINITY paper and Conductor paper — the ICLR 2026 papers behind the architecture
- LLM Gateway September 2026 timeline — confirmed Fugu Ultra v2 release date Sep 10, added to LLM Gateway Sep 11
- Sakana AI on X — Fugu Max and Fugu Ultra v2 announcement
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.