Sonnet 5.5 at $2 and $10 per MTok direct vs a Sume agent turn
Anthropic lists Sonnet 5.5 at $2 input and $10 output per million tokens. A Sume agent turn is list x 1.25 plus a 5.5% fee: $2.6375 and $13.1875. A worked turn.

Anthropic's own page lists Sonnet 5.5 at **$2 per million input tokens and $10 per million output tokens**. On Sume, an agent turn on the same model is list x 1.25, then a 5.5% fee: **$2.6375 input and $13.1875 output**. That is the same 1.31875 multiplier as media jobs.
The four Anthropic models
Anthropic's page also shows the Opus and Haiku rows. Sume's list prices in the repo match them. Multiply by 1.25, then by 1.055, to get the agent-turn price.
| Model | List input / output per MTok | Sume input / output per MTok |
|---|---|---|
| Haiku 4.5 | $1 / $5 | $1.31875 / $6.59375 |
| Sonnet 5.5 | $2 / $10 | $2.6375 / $13.1875 |
| Opus 5.5 | $4 / $20 | $5.275 / $26.375 |
A worked turn
Suppose an agent turn reads 100,000 uncached input tokens and writes 5,000 output tokens on Sonnet 5.5.
Direct: 0.1 x $2 + 0.005 x $10 = $0.25. On Sume: $0.25 x 1.31875 = $0.3296875. The $0.0796875 gap is what the 25% and the fee add.
Why an agent turn carries the fee
The 5.5% platform fee applies to all customer spend, including LLM turns. The base is the post-discount bill after the 1.25 step. Older constants in the repo describe some LLM models as having no Sume margin. The operations doc says that was superseded, and it is the document to trust.
What you get for the gap
An agent turn on Sume has its tools attached: generation, the sandbox, the files, and billing on one balance. A direct API call gives you tokens and nothing else. If you only need tokens, call the vendor. If the turn is going to fire a video job, one balance and one usage row per step is the reason to pay the multiplier.
Cache prices follow the same rule
Anthropic lists a 5-minute cache write at $2.50 per million tokens and a cache read at $0.20 on Sonnet 5.5. The same multiplier applies on Sume: a read is $0.20 x 1.31875 = $0.2637500 per million, and a write is $2.50 x 1.31875 = $3.296875.
Cache reads stay cheap in relative terms, because the ratio between a read and a fresh input token is unchanged. Long agent turns that re-read a large prompt benefit exactly as they do direct.
What to compare
Compare cost per finished task, not per token. A task that needs three tool calls, a sandbox and a video job is one balance on Sume. Done direct, it is an LLM bill, a sandbox bill and a video bill. The multiplier is the price of that consolidation, and you decide whether it is worth it for the workload.
Sources
Related posts
More in Pricing
- Soundstripe single-use licenses at $49, $199, $399 vs AI music
Soundstripe sells single-use licenses from $49 to $1,249+. Compare that per-use model with $0.125 per generated track and decide which one fits your video ads.
- Speech-to-text price per hour: Scribe, MAI Streaming, Sume
ElevenLabs Scribe v2 is $0.22 per hour, Realtime $0.39, Microsoft MAI streaming $0.54 and Sume STT $0.60. Costs for 1, 10 and 100 hours of audio.
- Split a long script into Sume TTS requests and price it
Sume TTS takes 20,000 characters per request at $0.0475 per 1,000. A Python splitter cuts at sentence ends and prices each request: 99,000 characters is $4.70.
- Twelve 750-character match recaps a weekend: a year of TTS cost
Twelve recaps of 750 characters every weekend for 52 weeks is 468,000 characters: $7.02 on MAI Flash, $10.30 on MAI-Voice-2.1, $22.23 on Sume TTS.
Written by Sume