Opus 5.5 fast mode is $8/$40 per million; a Format run has no tier
Anthropic prices Opus 5.5 fast mode at twice standard. Sume's Format run body has a model field but no service-tier field, so you cannot request it there.

You cannot ask for Claude Opus 5.5 fast mode on a Sume Format run, because the create body has no field for it. Anthropic lists Opus 5.5 at $4 per million input tokens and $20 per million output, and fast mode at $8 and $40 (read 2026-10-04). Sume's model field picks which LLM orchestrates the run, and the docs describe no service-tier, speed or priority option beside it.
That is a fact about the request shape, not a promise about what runs underneath, so treat fast mode as unavailable through Sume until the docs add a field.
What does the Sume body accept?
The Create a run page lists the fields: instruction, input, attachments, model, generation_spend_cap_usd, webhook_url, and the structured-output pair. An unknown model id is a 400 invalid_request, and the receipt echoes the id that ran.
| Mode | Input per million tokens | Output per million tokens |
|---|---|---|
| Standard | $4 | $20 |
| Fast mode | $8 | $40 |
| Cache read | $0.20 | n/a |
| Cache write | $5 | n/a |
Does speed matter for a video run?
Rarely. Most of a run's wall-clock time is spent waiting on generation, not on the orchestrator's tokens, so a faster LLM shortens the planning step but not the render. If latency is the problem, look at the run's events_url to see where time went before you pay for a faster tier somewhere you can use it.
What should I do instead?
If you need fast mode, call Anthropic directly for that step and use Sume for generation.
- Plan the shot list with fast mode in your own client.
- Send the finished plan to a Format run as
input, which is treated as data. - Cap the generation with
generation_spend_cap_usd. - Record the echoed
modelon each receipt so cost comparisons stay honest.
Sources
Related posts
More in Agents
- Opus 5.5 resists injection better, but keep scraped text in input
Anthropic says Opus 5.5 is more resistant to prompt injection. For a video agent that reads web pages, still send them to a Sume Format run as input data.
- Opus 5.5 hands cyber tasks to Opus 4.8: log the model id
Anthropic says Opus 5.5 re-routes most cybersecurity tasks to Opus 4.8. Keep the model id that served each Sume Format run in your logs to explain odd results.
- Codex Cloud background tasks: three Sume guardrails to set first
OpenAI's Codex Cloud runs tasks in the background while you are away. Before one can call Sume, set read-only scope, idempotency keys and a hard spend cap.
- Cursor Projects coordinator fan-out: size waves from generation_limits
A Cursor coordinator that delegates to subagents can overrun a Sume workspace. Budget new in-flight jobs from generation_limits, not from wave_size_hint.
Written by Sume