$ cat wiki/models/fugu-ultra-v2.md
Fugu Ultra v2
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Sakana AI |
| Released | 2026-09-11 |
| Announced | 2026-09-11 |
| Context window | 1,000,000 |
| Pricing | $5/M input · $30/M output · $0.50/M cached input; $10 / $45 / $1.00 above 272K context |
| License | unknown |
| Availability | Sakana API (OpenAI-compatible), OpenRouter, NanoGPT, Kilo Gateway, LLM Gateway |
| Catalogue id | sakana/fugu-ultra-v2 |
License reads unknown because **nothing read names a license for any Fugu | |
| model**. No open weights are offered; that is a statement about distribution, not a | |
| license term | |
| (source). |
What it is
The capability tier of the same orchestration architecture as Fugu Max — "the same core orchestration architecture optimized for two distinct missions" (2 passes) — targeting "absolute highest capability on complex, multi-step tasks". The architecture, the 7B conductor, the TRINITY and Conductor research basis and the arXiv 2606.21228 technical report are recorded once on Fugu Max and not repeated here.
Interface details published for this tier (source):
| Property | Value | Passes |
|---|---|---|
| Max completion tokens | 128,000 | 1 |
| Input modalities | text, images, and files such as PDFs; returns text | 1 |
| Tool calling | tools and tool_choice | 1 |
| Structured outputs | JSON schema in response_format | 1 |
| Web search | billed at $10.00/1K calls | 1 |
Benchmarks
All figures are Sakana's own claims from the release post, reached through search extracts. No independent reproduction exists in anything read, and there is no Artificial Analysis page for any Fugu model (1 pass).
| Result | Score | Passes |
|---|---|---|
| Chartography — Fugu Ultra v2 | 48.3 | 2 |
| Chartography — Claude Opus 5 | 27.3 | 2 |
| Chartography — Claude Fable 5 | 29.5 | 1 |
| DeepSWE — Fugu Ultra v2 | 74.3 | 2 |
| The two comparators are Claude Opus 5 and Claude Fable 5, and | ||
| they are named in the row labels rather than linked inside a cell: a table cell | ||
| holding a model name before its figure parses the digits in the name, which is how | ||
| this page's first draft reported a contradiction that did not exist. |
Ordinal claims: best or joint-best on five of eight benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon (2 passes) — and among the top two systems on seven of the eight (1 pass).
SWEFish is Sakana's own internal benchmark (2 passes), so one of the five is an evaluation the vendor wrote, holds and scores. Recorded here for the same reason it is recorded on Fugu Max: Eval Harness Configuration tracks exactly this, and a five-of-eight count reads differently once one of the eight is in-house.
The pool is the claim, not the score
Sakana states that Fugu Ultra v2 reaches these scores without Fable 5, Fable 5.1 or GPT-6 Astra in its agent pool (2 passes), and that it "does not rely on individual proprietary frontier models to deliver frontier output" (1 pass). One summary of the v2 change put it as "a smaller pool, a higher score" (source).
This is a stronger claim than a benchmark win and it is the one worth reading carefully. A router beating a frontier model is unsurprising — it can call that model. A router beating three frontier models it has removed from its own pool is an assertion that the capability sits in the coordination rather than in any single network, which is the thesis of the TRINITY and Conductor papers stated as a product.
What would test it, and is absent. The pool that remains is never enumerated, so the claim cannot be checked against what is actually in it; no ablation is published against a v2 run with the three models restored; and Chartography and DeepSWE are the only two numeric results in anything read, both from the vendor. The comparison figures are for the removed models individually — no orchestrated baseline over the same remaining pool appears anywhere read, so what the conductor contributes over a simpler policy is unmeasured.
Compared To
- Fugu Max — same architecture, wider pool, $2/M input · $6/M output against this page's $5/M input · $30/M output.
- Claude Opus 5, Claude Fable 5 — the two named numeric comparators, both on Chartography only.
- Claude Fable 5.1, Astra — named as excluded from the pool, which is a different relation from a benchmark comparison and is why they appear here rather than in the table.
- A Fugu score is a system score. Model Routing records why that is not the same claim as a model score — the number characterises an orchestrator over a pool, and this page cannot say what is in the pool.
Conflicting Reports
Release date. Two search passes give 2026-09-11; the MarkTechPost write-up sits at a URL dated 2026-09-10. 2026-09-11 is adopted, as the two-pass figure and the date attached to Sakana's own post. Unresolved — no first-party read was possible (source).
Sources
- source — release capture, 2026-09-11, search-extract only with per-figure pass counts
- Sakana AI release post — not read;
sakana.aianswersEGRESS_BLOCKEDfrom the cloud sandbox