Opus 5.5 fast mode is $8/$40 per million; a Format run has no tier

Anthropic prices Opus 5.5 fast mode at twice standard. Sume's Format run body has a model field but no service-tier field, so you cannot request it there.

5 min readSume
All posts

You cannot ask for Claude Opus 5.5 fast mode on a Sume Format run, because the create body has no field for it. Anthropic lists Opus 5.5 at $4 per million input tokens and $20 per million output, and fast mode at $8 and $40 (read 2026-10-04). Sume's model field picks which LLM orchestrates the run, and the docs describe no service-tier, speed or priority option beside it.

That is a fact about the request shape, not a promise about what runs underneath, so treat fast mode as unavailable through Sume until the docs add a field.

What does the Sume body accept?

The Create a run page lists the fields: instruction, input, attachments, model, generation_spend_cap_usd, webhook_url, and the structured-output pair. An unknown model id is a 400 invalid_request, and the receipt echoes the id that ran.

Opus 5.5 prices from Anthropic's announcement, read 2026-10-04.
ModeInput per million tokensOutput per million tokens
Standard$4$20
Fast mode$8$40
Cache read$0.20n/a
Cache write$5n/a

Does speed matter for a video run?

Rarely. Most of a run's wall-clock time is spent waiting on generation, not on the orchestrator's tokens, so a faster LLM shortens the planning step but not the render. If latency is the problem, look at the run's events_url to see where time went before you pay for a faster tier somewhere you can use it.

What should I do instead?

If you need fast mode, call Anthropic directly for that step and use Sume for generation.

  • Plan the shot list with fast mode in your own client.
  • Send the finished plan to a Format run as input, which is treated as data.
  • Cap the generation with generation_spend_cap_usd.
  • Record the echoed model on each receipt so cost comparisons stay honest.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume