Cost accounting

Sakana Fugu orchestration tokens are extra work, not a footnote

OpenAI-shaped field names make it easy to assume every token_details entry is a breakdown of a total you already counted. That holds for cached_tokens, which sits inside input_tokens — but official Sakana pricing documentation makes the orchestration fields additional real usage, billed at the matching input, cached, or output rate.

1. Store the official object, not a guessed subset

The live pricing and models pages publish this shape for Ultra (and, on the Models page, for Cyber as well):

usage
{
  "input_tokens": 120,
  "output_tokens": 80,
  "total_tokens": 200,
  "input_tokens_details": {
    "cached_tokens": 0,
    "orchestration_input_tokens": 0,
    "orchestration_input_cached_tokens": 0
  },
  "output_tokens_details": {
    "orchestration_output_tokens": 0
  }
}

Official field meanings, with the subset relationships made explicit:

  • input_tokens: total tokens from the user input sent to the first model, including any cached tokens
  • cached_tokens: the part of input_tokens served from cache — a subset, not an additional amount
  • orchestration_input_tokens: total input tokens used for orchestration, including its cached portion
  • orchestration_input_cached_tokens: the part of orchestration_input_tokens served from cache — a subset of that orchestration input
  • output_tokens: tokens in the final output
  • orchestration_output_tokens: output tokens from orchestration, billed on top of output_tokens
  • total_tokens: total including orchestration

The two cached fields follow ordinary OpenAI convention: each is a subset of its own input total, so cached_tokens ≤ input_tokens and orchestration_input_cached_tokens ≤ orchestration_input_tokens. Only the orchestration fields are extra work outside input_tokens and output_tokens; the ordinary input total already contains its cached portion.

2. Subtract each cached subset first, then sum, then apply rates

Uncached input = (input_tokens − cached_tokens) + (orchestration_input_tokens − orchestration_input_cached_tokens)

Billable cached input = cached_tokens + orchestration_input_cached_tokens

Billable output = output_tokens + orchestration_output_tokens

Estimate = uncached input/1M × input rate + cached input/1M × cached rate + output/1M × output rate

Validate before you price: each cached count must be less than or equal to its own input total, and each ordinary input total must include its cached subset. Reject a pair where the cached count is larger. Pricing cached tokens as if they were full-rate input — or adding them on top of an input total that already contains them — is the single most common way this estimate drifts by double the cache.

Orchestration totals stay extra: they are added to the ordinary request counts, not folded into them. A different model or provider may expose a different envelope, so do not parse an unvalidated usage object with this formula.

Ultra’s published standard rates on September 28, 2026 are $5 input, $0.50 cached, and $30 output per million tokens. Above 272K context the published rates become $10 / $1.00 / $45. Apply the higher tier only when that request’s context crossed the threshold.

3. A worked Ultra example with the same rates the calculator uses

Suppose a month of Ultra requests whose individual contexts each stay inside the standard tier aggregate to the following raw totals: 1,000,000 user input of which 250,000 is cached, 2,000,000 orchestration input of which 500,000 is cached, 250,000 final output, and 500,000 orchestration output. Each cached count is a subset of its own input total. This is an aggregate of many short requests, not one 1M-token request that risks the long-context tier.

  • Uncached input = (1,000,000 − 250,000) + (2,000,000 − 500,000) = 2,250,000 → $11.25
  • Billable cached input = 250,000 + 500,000 = 750,000 → $0.375
  • Billable output = 250,000 + 500,000 = 750,000 → $22.50
  • Total = $34.125, displayed as $34.13

If you had ignored orchestration, you would have stored $11.375 and argued with finance later. The Ultra and Max calculator is that arithmetic with the fields labeled and the cached subsets subtracted. It still cannot price standard Fugu or Cyber.

4. Operational rules that prevent a second argument

  • Persist the raw usage object beside the computed currency amount and the rate version you applied.
  • Do not assume total_tokens is priced at one blended rate. Input, cached, and output have different prices.
  • Subtract each cached count from its own input total before applying the input rate, and store the uncached figure you used.
  • Reject any usage object where a cached count exceeds its input total instead of silently clamping it.
  • Do not treat Ultra max_output_tokens as a spend cap. Official docs say it limits the final model, not the orchestrator.
  • Split mixed traffic: requests under 272K and requests over 272K are two calculations.
  • If a partner route does not return these fields, you cannot reuse this formula without new evidence.

Keep the raw JSON even when the computed dollar amount looks tidy. A later pricing-page change, or a later realization that cached orchestration was omitted, is cheaper to fix from stored fields than from a screenshot of a total.

Sources