Opus 5.5 costs 20% less than Opus 5: video agent turn math

Opus 5.5 lists $4 input and $20 output per MTok against Opus 5 at $5 and $25. Priced on a 200K-token video-agent turn, and what Sume did with Opus 5.

5 min readSume
All posts

Claude Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Claude Opus 5, which is 20 percent less on both sides. For a video-agent turn that sends 200K tokens of thread and writes 10K tokens back, that is $1.00 on Opus 5.5 instead of $1.25.

Prices, context and output limits come from the Claude Platform release notes, entry of September 22, 2026, read 2026-10-03. The same entry says Opus 5.5 has a 1M-token context window by default, 128k max output tokens and always-on adaptive thinking. Sume facts come from the model catalog source.

What does a turn cost at these list prices?

The arithmetic is list price times tokens, ignoring caching, which changes the input side. Output includes any thinking the model produces, since the notes describe thinking as always on and controlled by effort; so a turn's output can exceed the visible text. Treat the output figure as a floor.

Opus 5 vs Opus 5.5 list cost per turn (read 2026-10-03)
Turn shapeOpus 5 ($5 / $25)Opus 5.5 ($4 / $20)Saving
50K in, 2K out$0.30$0.24$0.06
200K in, 10K out$1.25$1.00$0.25
800K in, 20K out$4.50$3.60$0.90
1M in, 50K out (full window plus long answer)$6.25$5.00$1.25

Why does a video agent care about the input side?

A video agent's thread grows with every tool result: job ids, probe facts, shot lists, review notes. The input side is re-sent each turn unless it is cached, so a thread that reaches 500K tokens pays for 500K on every later turn. At $4 per million that is $2.00 per turn on Opus 5.5 against $2.50 on Opus 5. The saving compounds with thread length, which is the argument for ending a thread and starting a new one with a short summary instead of letting it run to the window.

The 1M default window removes the hard stop, not the cost. Long threads are allowed; they are just not free.

What did Sume do with Opus 5?

In the agent model catalog, Opus 5.5 is listed (openrouter/anthropic-claude-opus-5-5) and described as built for long-running agentic coding and knowledge work, 1M context and 128K max output. Opus 5 is marked retired onto Opus 5.5: it is never listed for new picks, kept only so persisted rows keep their exact id and their original pricing card. In other words, a stored Opus 5 thread still prints and prices as Opus 5, while a new pick is Opus 5.5.

For API callers, the model field on Calling a Format takes an Agents catalog id, and an id outside the catalog is 400 invalid_request.

What should I check before switching a pin?

The notes list API changes on Opus 5.5 that affect direct integrations: thinking: {"type": "disabled"} and thinking: {"type": "enabled", ...} both return 400, so omit thinking and use effort; tool_choice types any and tool return 400; and computer use needs computer_toolset_20260801 on the Claude API and Google Cloud. A cheaper model that rejects your request shape is not cheaper until you change the request.

Through Sume none of those fields are yours to set. The OpenRouter-routed rows ship no effort or thinking parameters, so the swap is a single model id in the request.

curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: opus-55-pin-001" \
  -d '{"instruction":"Three hook ideas for a vitamin C serum",
       "model":"openrouter/anthropic-claude-opus-5-5",
       "generation_spend_cap_usd":5}'

Does a cheaper Opus change which model I pick?

It narrows the gap, but it does not change the shape of the decision. Pricing tells you what a turn costs; it does not tell you whether the extra capability is needed for that turn. Planning a twelve-shot ad with strict rules is a plausible Opus job; labelling frames or rewriting a caption is not. A mixed setup is common: a stronger orchestrator for the plan, a cheaper model for repeated checks.

On Sume the orchestrator is one model per run, set with model, so a mixed setup means separate runs rather than one thread that swaps. That is also the arrangement that avoids the thread-switching problems described in the release notes for thinking blocks.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume