Opus 5.5 costs 20% less than Opus 5: video agent turn math
Opus 5.5 lists $4 input and $20 output per MTok against Opus 5 at $5 and $25. Priced on a 200K-token video-agent turn, and what Sume did with Opus 5.

Claude Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Claude Opus 5, which is 20 percent less on both sides. For a video-agent turn that sends 200K tokens of thread and writes 10K tokens back, that is $1.00 on Opus 5.5 instead of $1.25.
Prices, context and output limits come from the Claude Platform release notes, entry of September 22, 2026, read 2026-10-03. The same entry says Opus 5.5 has a 1M-token context window by default, 128k max output tokens and always-on adaptive thinking. Sume facts come from the model catalog source.
What does a turn cost at these list prices?
The arithmetic is list price times tokens, ignoring caching, which changes the input side. Output includes any thinking the model produces, since the notes describe thinking as always on and controlled by effort; so a turn's output can exceed the visible text. Treat the output figure as a floor.
| Turn shape | Opus 5 ($5 / $25) | Opus 5.5 ($4 / $20) | Saving |
|---|---|---|---|
| 50K in, 2K out | $0.30 | $0.24 | $0.06 |
| 200K in, 10K out | $1.25 | $1.00 | $0.25 |
| 800K in, 20K out | $4.50 | $3.60 | $0.90 |
| 1M in, 50K out (full window plus long answer) | $6.25 | $5.00 | $1.25 |
Why does a video agent care about the input side?
A video agent's thread grows with every tool result: job ids, probe facts, shot lists, review notes. The input side is re-sent each turn unless it is cached, so a thread that reaches 500K tokens pays for 500K on every later turn. At $4 per million that is $2.00 per turn on Opus 5.5 against $2.50 on Opus 5. The saving compounds with thread length, which is the argument for ending a thread and starting a new one with a short summary instead of letting it run to the window.
The 1M default window removes the hard stop, not the cost. Long threads are allowed; they are just not free.
What did Sume do with Opus 5?
In the agent model catalog, Opus 5.5 is listed (openrouter/anthropic-claude-opus-5-5) and described as built for long-running agentic coding and knowledge work, 1M context and 128K max output. Opus 5 is marked retired onto Opus 5.5: it is never listed for new picks, kept only so persisted rows keep their exact id and their original pricing card. In other words, a stored Opus 5 thread still prints and prices as Opus 5, while a new pick is Opus 5.5.
For API callers, the model field on Calling a Format takes an Agents catalog id, and an id outside the catalog is 400 invalid_request.
What should I check before switching a pin?
The notes list API changes on Opus 5.5 that affect direct integrations: thinking: {"type": "disabled"} and thinking: {"type": "enabled", ...} both return 400, so omit thinking and use effort; tool_choice types any and tool return 400; and computer use needs computer_toolset_20260801 on the Claude API and Google Cloud. A cheaper model that rejects your request shape is not cheaper until you change the request.
Through Sume none of those fields are yours to set. The OpenRouter-routed rows ship no effort or thinking parameters, so the swap is a single model id in the request.
curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: opus-55-pin-001" \
-d '{"instruction":"Three hook ideas for a vitamin C serum",
"model":"openrouter/anthropic-claude-opus-5-5",
"generation_spend_cap_usd":5}'Does a cheaper Opus change which model I pick?
It narrows the gap, but it does not change the shape of the decision. Pricing tells you what a turn costs; it does not tell you whether the extra capability is needed for that turn. Planning a twelve-shot ad with strict rules is a plausible Opus job; labelling frames or rewriting a caption is not. A mixed setup is common: a stronger orchestrator for the plan, a cheaper model for repeated checks.
On Sume the orchestrator is one model per run, set with model, so a mixed setup means separate runs rather than one thread that swaps. That is also the arrangement that avoids the thread-switching problems described in the release notes for thinking blocks.
Sources
Related posts
More in Pricing
- Cost per row of a Sume bulk run: add up debited_usd_micros
Read usage.debited_usd_micros on each child receipt, wait for final to be true, and treat null as unknown. Why billable_amount alone understates a bulk run.
- Cost to remove backgrounds and upscale 500 product photos by API
Sume bills background removal at $0.0225 an image and upscale at $0.20, both flat. 500 photos through both steps cost $111.25; the table and a submit loop.
- Does AI video API billing round up? Sume's ceil rules per endpoint
A 5.2 s clip is billed as 6 s. Which Sume endpoints round seconds or minutes up, which prorate, and worked examples from the published rates.
- Eleven v4 from $6 a month vs Sume's $0.0475 per 1,000 characters
ElevenLabs' v4 page lists a free tier of 10,000 credits and paid plans from $6 a month; Sume bills TTS per character. Which shape fits an occasional voiceover.
Written by Sume