Cost accounting
Sakana Fugu orchestration tokens are extra work, not a footnote
OpenAI-shaped field names make it easy to assume every token_details entry is a breakdown of a total you already counted. That holds for cached_tokens, which sits inside input_tokens — but official Sakana pricing documentation makes the orchestration fields additional real usage, billed at the matching input, cached, or output rate.
1. Store the official object, not a guessed subset
The live pricing and models pages publish this shape for Ultra (and, on the Models page, for Cyber as well):
{
"input_tokens": 120,
"output_tokens": 80,
"total_tokens": 200,
"input_tokens_details": {
"cached_tokens": 0,
"orchestration_input_tokens": 0,
"orchestration_input_cached_tokens": 0
},
"output_tokens_details": {
"orchestration_output_tokens": 0
}
}
Official field meanings, with the subset relationships made explicit:
input_tokens: total tokens from the user input sent to the first model, including any cached tokenscached_tokens: the part ofinput_tokensserved from cache — a subset, not an additional amountorchestration_input_tokens: total input tokens used for orchestration, including its cached portionorchestration_input_cached_tokens: the part oforchestration_input_tokensserved from cache — a subset of that orchestration inputoutput_tokens: tokens in the final outputorchestration_output_tokens: output tokens from orchestration, billed on top ofoutput_tokenstotal_tokens: total including orchestration
The two cached fields follow ordinary OpenAI convention: each is a subset of its own input total, so cached_tokens ≤ input_tokens and orchestration_input_cached_tokens ≤ orchestration_input_tokens. Only the orchestration fields are extra work outside input_tokens and output_tokens; the ordinary input total already contains its cached portion.
2. Subtract each cached subset first, then sum, then apply rates
Uncached input = (input_tokens − cached_tokens) + (orchestration_input_tokens − orchestration_input_cached_tokens)
Billable cached input = cached_tokens + orchestration_input_cached_tokens
Billable output = output_tokens + orchestration_output_tokens
Estimate = uncached input/1M × input rate + cached input/1M × cached rate + output/1M × output rate
Validate before you price: each cached count must be less than or equal to its own input total, and each ordinary input total must include its cached subset. Reject a pair where the cached count is larger. Pricing cached tokens as if they were full-rate input — or adding them on top of an input total that already contains them — is the single most common way this estimate drifts by double the cache.
Orchestration totals stay extra: they are added to the ordinary request counts, not folded into them. A different model or provider may expose a different envelope, so do not parse an unvalidated usage object with this formula.
Ultra’s published standard rates on September 28, 2026 are $5 input, $0.50 cached, and $30 output per million tokens. Above 272K context the published rates become $10 / $1.00 / $45. Apply the higher tier only when that request’s context crossed the threshold.
3. A worked Ultra example with the same rates the calculator uses
Suppose a month of Ultra requests whose individual contexts each stay inside the standard tier aggregate to the following raw totals: 1,000,000 user input of which 250,000 is cached, 2,000,000 orchestration input of which 500,000 is cached, 250,000 final output, and 500,000 orchestration output. Each cached count is a subset of its own input total. This is an aggregate of many short requests, not one 1M-token request that risks the long-context tier.
- Uncached input = (1,000,000 − 250,000) + (2,000,000 − 500,000) = 2,250,000 → $11.25
- Billable cached input = 250,000 + 500,000 = 750,000 → $0.375
- Billable output = 250,000 + 500,000 = 750,000 → $22.50
- Total = $34.125, displayed as $34.13
If you had ignored orchestration, you would have stored $11.375 and argued with finance later. The Ultra and Max calculator is that arithmetic with the fields labeled and the cached subsets subtracted. It still cannot price standard Fugu or Cyber.
4. Operational rules that prevent a second argument
- Persist the raw usage object beside the computed currency amount and the rate version you applied.
- Do not assume
total_tokensis priced at one blended rate. Input, cached, and output have different prices. - Subtract each cached count from its own input total before applying the input rate, and store the uncached figure you used.
- Reject any usage object where a cached count exceeds its input total instead of silently clamping it.
- Do not treat Ultra
max_output_tokensas a spend cap. Official docs say it limits the final model, not the orchestrator. - Split mixed traffic: requests under 272K and requests over 272K are two calculations.
- If a partner route does not return these fields, you cannot reuse this formula without new evidence.
Keep the raw JSON even when the computed dollar amount looks tidy. A later pricing-page change, or a later realization that cached orchestration was omitted, is cheaper to fix from stored fields than from a screenshot of a total.