AI Trend Notifier
EN
← wiki

$ cat wiki/models/fugu-ultra-v2.md

Fugu Ultra v2

modelupdated 2026-09-12created 2026-09-12

Compared with

Spec

AttributeValue
DeveloperSakana AI
Released2026-09-11
Announced2026-09-11
Context window1,000,000
Pricing$5/M input · $30/M output · $0.50/M cached input; $10 / $45 / $1.00 above 272K context
Licenseunknown
AvailabilitySakana API (OpenAI-compatible), OpenRouter, NanoGPT, Kilo Gateway, LLM Gateway
Catalogue idsakana/fugu-ultra-v2
License reads unknown because **nothing read names a license for any Fugu
model**. No open weights are offered; that is a statement about distribution, not a
license term
(source).

Release Date

2026-09-11, alongside Fugu Max, in Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier (2 passes) (source). One source dates it 2026-09-10; see ## Conflicting Reports.

What it is

The capability tier of the same orchestration architecture as Fugu Max"the same core orchestration architecture optimized for two distinct missions" (2 passes) — targeting "absolute highest capability on complex, multi-step tasks". The architecture, the 7B conductor, the TRINITY and Conductor research basis and the arXiv 2606.21228 technical report are recorded once on Fugu Max and not repeated here.

Interface details published for this tier (source):

PropertyValuePasses
Max completion tokens128,0001
Input modalitiestext, images, and files such as PDFs; returns text1
Tool callingtools and tool_choice1
Structured outputsJSON schema in response_format1
Web searchbilled at $10.00/1K calls1

Benchmarks

All figures are Sakana's own claims from the release post, reached through search extracts. No independent reproduction exists in anything read, and there is no Artificial Analysis page for any Fugu model (1 pass).

ResultScorePasses
Chartography — Fugu Ultra v248.32
Chartography — Claude Opus 527.32
Chartography — Claude Fable 529.51
DeepSWE — Fugu Ultra v274.32
The two comparators are Claude Opus 5 and Claude Fable 5, and
they are named in the row labels rather than linked inside a cell: a table cell
holding a model name before its figure parses the digits in the name, which is how
this page's first draft reported a contradiction that did not exist.

Ordinal claims: best or joint-best on five of eight benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon (2 passes) — and among the top two systems on seven of the eight (1 pass).

SWEFish is Sakana's own internal benchmark (2 passes), so one of the five is an evaluation the vendor wrote, holds and scores. Recorded here for the same reason it is recorded on Fugu Max: Eval Harness Configuration tracks exactly this, and a five-of-eight count reads differently once one of the eight is in-house.

The pool is the claim, not the score

Sakana states that Fugu Ultra v2 reaches these scores without Fable 5, Fable 5.1 or GPT-6 Astra in its agent pool (2 passes), and that it "does not rely on individual proprietary frontier models to deliver frontier output" (1 pass). One summary of the v2 change put it as "a smaller pool, a higher score" (source).

This is a stronger claim than a benchmark win and it is the one worth reading carefully. A router beating a frontier model is unsurprising — it can call that model. A router beating three frontier models it has removed from its own pool is an assertion that the capability sits in the coordination rather than in any single network, which is the thesis of the TRINITY and Conductor papers stated as a product.

What would test it, and is absent. The pool that remains is never enumerated, so the claim cannot be checked against what is actually in it; no ablation is published against a v2 run with the three models restored; and Chartography and DeepSWE are the only two numeric results in anything read, both from the vendor. The comparison figures are for the removed models individually — no orchestrated baseline over the same remaining pool appears anywhere read, so what the conductor contributes over a simpler policy is unmeasured.

Use Cases

Stated only as a mission: complex, multi-step tasks, against Fugu Max's cost-performance target (2 passes) (source). The published web-search billing line implies tool-using agent workloads; nothing read names a deployment.

Compared To

  • Fugu Max — same architecture, wider pool, $2/M input · $6/M output against this page's $5/M input · $30/M output.
  • Claude Opus 5, Claude Fable 5 — the two named numeric comparators, both on Chartography only.
  • Claude Fable 5.1, Astra — named as excluded from the pool, which is a different relation from a benchmark comparison and is why they appear here rather than in the table.
  • A Fugu score is a system score. Model Routing records why that is not the same claim as a model score — the number characterises an orchestrator over a pool, and this page cannot say what is in the pool.

Conflicting Reports

Release date. Two search passes give 2026-09-11; the MarkTechPost write-up sits at a URL dated 2026-09-10. 2026-09-11 is adopted, as the two-pass figure and the date attached to Sakana's own post. Unresolved — no first-party read was possible (source).

Sources

  • source — release capture, 2026-09-11, search-extract only with per-figure pass counts
  • Sakana AI release post — not read; sakana.ai answers EGRESS_BLOCKED from the cloud sandbox

Referenced by

Sources