Model guide

Fugu Ultra: what it is, what it costs, and when it beats Fugu Max

Short answer: Fugu Ultra is the quality-first model in Sakana AI's Sakana Fugu family. It coordinates a deeper, fixed pool of expert agents for hard multi-step work and answers more slowly than the default Fugu. The current version, fugu-ultra-v2.0, costs $5 per million input tokens and $30 per million output tokens up to 272K context, plus orchestration tokens, and it is included in every $20, $100 and $200 subscription tier. Use it when one wrong answer costs more than the tokens; use Fugu Max when cost per task matters most. You can price a real request below.

1. What Fugu Ultra is

Sakana Fugu is not a single large language model. It is an orchestration system sold as a model: one OpenAI-compatible API call goes in, and Fugu decides which underlying models to involve, which roles they play and how their work is combined. Sakana AI sells four variants under that name. The default fugu balances quality with low latency, fugu-max is tuned for cost-performance, fugu-cyber is specialized for security work behind an access request, and fugu-ultra is the variant tuned for output quality.

The official product page describes Ultra as coordinating "a deeper pool of expert agents to maximize answer quality on hard, high-stakes problems." Its FAQ adds the trade-off in plain terms: Ultra coordinates more expert agents when accuracy and depth matter most, at the cost of response time. Sakana lists early uses such as Kaggle competitions, paper reproduction, cybersecurity analysis, and literature and patent investigations. Those are long, multi-step jobs where a model is allowed to think for a while, not chat replies that must appear in a second.

Two properties set Ultra apart from the default model. First, its agent pool is fixed: Sakana says Ultra relies on the full pool to reach its performance, so you cannot opt individual providers or models out of it the way you can with fugu. Second, the routing is proprietary. The API reports token usage and cost per request, but it does not reveal which underlying models answered a given query. If your compliance team needs a list of every model that touched the data, Ultra cannot give it to you.

The version matters too. The moving fugu-ultra alias currently points to v2.0, the official page states that Ultra v2.0's training cutoff is 2026-08-28, and it names Fable 5, Fable 5.1 and GPT-6-Astra as models that are not in Ultra's pool. Sakana also says that when a new frontier model is released it expects to spend roughly two weeks training and evaluating updated Fugu models before rollout. If you need reproducible results across that kind of change, read when to pin an Ultra version.

2. What Sakana claims, and what it does not

Sakana's product page says Fugu Ultra v2.0 reaches the best or joint-best score on five of eight benchmarks it reports (GDP.pdf, Chartography, DeepSWE, Toolathon and SWEFish, an internal Sakana coding benchmark) and places in the top two on seven of the eight. It also publishes qualitative case studies, such as an AutoResearch run where Ultra finished with the best mean validation bits-per-byte against three anonymized frontier baselines, and a Rubik's Cube solver task where its program solved all 300 held-out cubes.

These are vendor-reported results. This site has not reproduced them, one of the eight benchmarks is Sakana's own, and the baselines in the case studies are deliberately anonymized and reshuffled between examples. Treat the numbers as a reason to test Ultra, not as proof that it will beat the model you use today on your workload. The benchmark evaluation playbook shows how to build a small, fair comparison with your own prompts, a fixed scoring rule and a cost column.

3. Fugu Ultra price and billing

ItemUp to 272K request contextAbove 272K
Input, per 1M tokens$5.00$10.00
Cached input, per 1M tokens$0.50$1.00
Output, per 1M tokens$30.00$45.00
Orchestration tokensBilled on top, at the same input, cached and output rates

Those are pay-as-you-go rates for fugu-ultra-v2.0. Pay-as-you-go tokens are served at higher priority than monthly-plan tokens, which is why Sakana recommends the token plan for heavy production work. The alternative is a subscription: Standard at $20 a month, Pro at $100 (ten times Standard usage) and Max at $200 (twenty times Standard usage). Every tier includes Fugu, Fugu Ultra and Fugu Max, but the size of the Standard allowance in tokens is not published, so you cannot work out in advance how many Ultra requests a subscription covers. The trade-offs are compared in subscription versus pay-as-you-go.

The line that catches people out is orchestration. Ultra does extra work between agents, and the usage object reports it in separate fields: orchestration_input_tokens, orchestration_input_cached_tokens and orchestration_output_tokens. Those sit outside the ordinary input_tokens and output_tokens totals. An application that logs only the ordinary totals will under-report Ultra spend. The orchestration token guide has the full formula.

A worked example: a code review request that sends 120,000 input tokens, 20,000 of them cached, and receives 6,000 output tokens. At Ultra's base rates the visible part costs $0.50 for the 100,000 uncached input tokens, $0.01 for the cached part and $0.18 for the output, so about $0.69 before orchestration. The same visible usage on Fugu Max costs $0.20 + $0.005 + $0.036, about $0.24. Push the same request past 272K of context and Ultra's input rate doubles, which is why trimming the prompt matters more on Ultra than on Max.

4. Fugu Ultra vs Fugu Max vs Fugu

QuestionFugu UltraFugu MaxFugu (default)
Optimized forOutput quality on hard, multi-step workCost-performanceBalanced quality and low latency
SpeedSlower; official FAQ trades response time for qualityNot stated as a speed modelPositioned for interactive chat and coding
Pay-as-you-go rate$5 / $0.50 / $30 per 1M up to 272K; $10 / $1.00 / $45 aboveFlat $2 / $0.25 / $6 per 1M at any context lengthStandard rate of the underlying model; never stacked
Extra billed usageOrchestration tokens$0.007 per web_search or web_fetch callSingle rate of the top-tier model involved
Can you opt models out?No, fixed poolNo, fixed poolYes, in the console settings
Included in subscriptionsYes, every tierYes, every tierYes, every tier

A naming trap: the $200 Max plan is a subscription tier, while Fugu Max is a model. Paying for the Max plan does not mean your requests run on Fugu Max, and Ultra is available on the $20 tier as well. For a guided pick based on your task, latency and budget, use the model chooser.

5. Calling Fugu Ultra correctly

  • Use the Responses API for new work. It accepts OpenAI-style requests, so most clients only need a new base URL, key and model ID. The API quickstart covers setup.
  • Resend history yourself. previous_response_id is not accepted, so multi-turn work sends the whole conversation each time. On Ultra that history is billed at $5 per million input tokens, or $0.50 when cached.
  • Do not treat max_output_tokens as a spend cap. On Ultra it limits the final answer only, not the orchestrator's own work.
  • Log the orchestration fields. Store all three orchestration counts per request so finance sees the real cost.
  • Set timeouts for long jobs. Ultra is designed to take longer. Run it in background jobs or queues rather than inside a request a user is waiting on.
  • Check the region first. Sakana Fugu is not offered in the EU or EEA; see EU availability.

6. Price your own Ultra request

Paste the usage object from one real Ultra response. Enter the full input total including cached tokens, then the cached part once, and add the orchestration fields that Ultra reports separately. Switch the model to Fugu Max to compare the same visible usage at Max rates. This is the same published-rate calculator as the calculator page; Cyber rates are not public and the default Fugu model is billed by route, so neither is included.

Fugu Ultra charges one rate up to 272K request context and a higher rate above it.

User-visible usage

Enter the full input_tokens total from the API response: it already counts cached_tokens. List the cached part once as well, because the calculator bills the rest at the input rate and the cached part at the cached rate. Do not add cached tokens on top of the total.

Orchestration usage

These fields are extra Ultra usage outside the ordinary request totals above, not a breakdown of them. orchestration_input_tokens likewise includes orchestration_input_cached_tokens. Leave them at zero to cost only the visible request and response.

Tool calls

Fugu Max publishes $0.007 per web_search or web_fetch call. No per-call rate is published for Ultra, so this field must stay at zero for Ultra.

7. Should you use Fugu Ultra?

Use Ultra for expensive mistakes

Code review before a release, a literature or patent survey, a security analysis or a paper reproduction: jobs where missing one issue costs more than a few dollars of tokens and nobody is waiting on the reply.

Use Max for volume

Batch jobs, summaries and agent loops that run thousands of times. Max's flat $2 / $6 rate and its cost-performance tuning make the bill far easier to forecast.

Use the default Fugu for interactive work

Chat interfaces and coding assistants where latency is visible to the user, or when you must opt specific providers out for compliance.

Do not use Ultra if

you need to name every model that processed the data, your users are in the EU or EEA, or you cannot log orchestration usage. Fix those first.

Frequently asked questions

What is Fugu Ultra?

Fugu Ultra is the quality-first model in Sakana AI's Sakana Fugu family. It orchestrates a deeper, fixed pool of expert agents to maximize answer quality on hard multi-step problems, at the cost of longer response times. The current version is fugu-ultra-v2.0.

How much does Fugu Ultra cost?

On pay-as-you-go, Fugu Ultra v2.0 costs $5 per million input tokens, $0.50 cached input and $30 output up to 272K request context, and $10 / $1.00 / $45 above 272K. Orchestration tokens are billed on top at the same category rates. Every $20, $100 and $200 monthly subscription tier also includes Ultra.

Is Fugu Ultra better than Fugu Max?

Not for every job. Sakana positions Ultra for raw output quality and Max for cost-performance. Max is a flat $2 input and $6 output per million tokens, so its output is one fifth of Ultra's base output rate. Test both on your own tasks before choosing.

Can I choose which models Fugu Ultra uses?

No. Sakana's FAQ says Ultra relies on its full agent pool, so the pool is fixed, and the routing for each query is not shown. Opting specific models out is only offered for the default Fugu model.

Is Fugu Ultra free or available in the EU?

There is no permanent free tier for the API; the cheapest way in is the $20 Standard subscription. Sakana Fugu is not offered to users in the EU or EEA, and this site does not recommend routing around that restriction.

Sources and verification boundary

SakanaFugu.com is an independent guide and is not affiliated with Sakana AI. Benchmark results above are Sakana's own and were not reproduced here. Recheck the official pricing page before you commit a budget.