$ cat wiki/models/fugu-max.md
Fugu Max
Compared with
- DeepSeek V4.1-Flash
- GPT-Image-2.5 Flare
- K2 Horizon
- Gemini 3.8 Flash
- Muse Spark 1.3
- Claude Fable 5.1
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Grok 4.6
- Laguna S 2.1
- Kimi K3
- Inkling
- LongCat-2.0
- MiniMax M3
- Fugu Ultra v2
- GPT-Image-2.5 Sunburst
- Astra
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
Spec
| Attribute | Value |
|---|---|
| Developer | Sakana AI |
| Released | 2026-09-11 |
| Announced | 2026-09-11 |
| Context window | 1,000,000 |
| Pricing | $2/M input · $6/M output |
| License | unknown |
| Availability | Sakana API (OpenAI-compatible), OpenRouter, NanoGPT, Kilo Gateway, LLM Gateway |
| Catalogue id | sakana/fugu-max |
License reads unknown rather than proprietary because **nothing read names a | |
| license for any Fugu model**. What is stated is that no open weights are offered | |
| — the product is the endpoint — and "no weights" is not the same statement as a | |
| license term | |
| (source). |
The Catalogue id row is present so the daily spec-check Action can resolve one:
the slug fugu-max does not reach the catalogue path, which is namespaced under
the vendor. Both ids were copied from catalogue URLs returned in search results,
not constructed.
Release Date
2026-09-11, alongside Fugu Ultra v2, in a post titled Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier (2 passes) (source).
One source dates it 2026-09-10 — the MarkTechPost write-up's URL path. Recorded
in ## Conflicting Reports below rather than resolved.
What it is — a model whose parameters belong to other labs
Fugu Max is not a single trained network. It is a learned multi-agent orchestrator: a router that hands each task to a pool of other models and stitches the answers back into one response, served behind a single OpenAI-compatible endpoint. Sakana's phrase is "a Multi-Agent System, Delivered as One Model", reaching its results "by dynamically coordinating and orchestrating a diverse pool of powerful models" (2 passes) (source).
| Property | Value | Passes |
|---|---|---|
| Conductor | a trained 7B-parameter model that learns which models to activate, how agents communicate, and how to combine their work | 1 |
| Research basis | two ICLR 2026 papers, TRINITY and Conductor | 2 |
| TRINITY | an evolved coordinator assigning Thinker, Worker or Verifier roles across several turns | 1 |
| Conductor | trained with reinforcement learning to discover natural-language coordination strategies and prompts | 2 |
| Self-call | the router can recursively call instances of itself | 1 |
| Technical report | arXiv 2606.21228 | 2 |
| **What separates Max from Fugu Ultra v2 is the pool and the mission, not | ||
| the architecture.** Sakana describes the two as *"the same core orchestration | ||
| architecture optimized for two distinct missions"* (2 passes). Max widens the pool to | ||
| "an unprecedented number of open-weights and specialized models, including NVIDIA Nemotron family through our collaboration with NVIDIA" | ||
| (2 passes) — see Nemotron 3.5 Lightning and NVIDIA. |
The pool is never enumerated in anything read. It is described by category and by one named family. A benchmark score produced by an unlisted pool cannot be attributed to any component of it.
Benchmarks
All figures below are Sakana's own claims from the release post, reached through search extracts. No independent reproduction exists in anything read, and there is no Artificial Analysis page for any Fugu model (1 pass) (source).
| Claim | Value | Passes |
|---|---|---|
| Best overall score | six benchmarks — Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, SWEFish | 2 |
| Cost-performance Pareto frontier | expanded on 7 of 10 benchmarks | 2 |
| SWEFish is Sakana's own benchmark. The release post describes it as *"our | ||
| internal benchmark reflecting Sakana AI's own coding challenges and use-cases"* | ||
| (2 passes). One of the six "best overall" results is therefore on an evaluation the | ||
| vendor wrote, holds and scores — recorded here because | ||
| Eval Harness Configuration is the page that tracks exactly this, and | ||
| because the six-benchmark count reads differently once one of the six is in-house. |
No per-benchmark score is published in anything read for Fugu Max — the claims are ordinal ("best overall", "expands the frontier"), not numeric. Fugu Ultra v2 has two named scores; Max has none.
Pricing
$2/M input tokens · $6/M output tokens (2 passes). Sakana's comparison claim: output pricing 40–60% lower than Sonnet 5, GPT 5.6 Terra and Kimi K3 (2 passes).
GPT 5.6 Terra has no page on this wiki and is named here as the release post names it, without a link. This wiki holds GPT-5.6 Sol (and Terra, Luna) and GPT-5.6-Cyber; nothing read connects Terra to either.
The published rate is what the caller pays, not what the request costs. Fugu Max's answer is assembled from calls to other vendors' models at those vendors' rates, and nothing read decomposes a request, states how many downstream calls one makes, or says who absorbs the difference between the two figures. This is the gap Model Routing names in its own words — "the published price of a model stops describing what a workflow costs" — arriving as a published price on a model page for the first time.
Use Cases
Stated by Sakana only as a mission rather than a workload list: Max answers "What is the best possible output we can deliver at the lowest possible cost?", against Ultra v2's target of complex multi-step work (1 pass) (source).
Its stated thesis is about open models: "Open models become dramatically more useful when orchestrated together rather than used in isolation", and Max is "a bet that the Pareto frontier of the future will be built out of many open, specialized models working in concert" (1 pass). That is a position on Open-Weights Policy Fight expressed as a product tier.
Compared To
- Fugu Ultra v2 — the capability tier of the same architecture, at $5/M input · $30/M output against Max's $2/M input · $6/M output. Same architecture, same context window, different pool and different price.
- Nemotron 3.5 Lightning — shipped the same day as NVIDIA's NeMo Switchyard router. NVIDIA shipped a cheap model and the software that decides when to use one, as two products; Sakana sells the deciding as the product and puts Nemotron inside it. The comparison is recorded on Model Routing.
- Claude Sonnet 5, Kimi K3 — the two named price comparators with pages here.
- The models it orchestrates are not comparators. A Fugu score is a system score, and Model Routing records why that is a different claim.
Conflicting Reports
Release date. Two search passes give 2026-09-11, one of them stating it as the capture-day date; the MarkTechPost write-up carrying the launch sits at a URL dated 2026-09-10. 2026-09-11 is adopted as the two-pass figure and as the date attached to Sakana's own post. Unresolved — no first-party read was possible (source).
Sources
- source — release capture, 2026-09-11, search-extract only with per-figure pass counts
- Sakana AI release post — not read;
sakana.aianswersEGRESS_BLOCKEDfrom the cloud sandbox