Sakana AI announced the release of Fugu Max and Fugu Ultra v2 on September 11, 2026. Both models use the company's shared orchestration architecture, with Fugu Max targeted at cost-performance and Ultra v2 aimed at complex tasks. Both are available through its OpenAI-compatible API in supported regions.
The pools available to Max and Ultra v2 customers are fixed. While standard Fugu permits individual model opt-outs, Sakana's product documentation states that Max and Ultra v2 do not support user opt-outs, with customized configurations offered separately to enterprise customers. Sakana states that Fugu Max routes across open and specialized models, including NVIDIA Nemotron, while Fugu Ultra v2's pool excludes Fable 5, Fable 5.1, and GPT-6 Astra, with an August 28, 2026 training cutoff.
The product FAQ says Ultra prioritizes answer quality at the cost of response time. Sakana’s performance positioning and benchmark results are company claims, rather than independent tests reported here.
Pricing is split by tier and context window, billed per million tokens:
| Model / Feature | Input (per 1M) | Output (per 1M) | Cached Input (per 1M) |
|---|---|---|---|
| Fugu Max (all contexts) | $2.00 | $6.00 | $0.25 |
| Fugu Ultra v2 (≤272K context) | $5.00 | $30.00 | $0.50 |
| Fugu Ultra v2 (>272K context) | $10.00 | $45.00 | $1.00 |
Tool calls using web_search or web_fetch incur an additional $0.007 fee per call, in addition to token charges.
