A 20-turn agent session: Haiku 5.5 $0.03, Astra $3.10 with caching
Cache write on turn one, reads after: a 20-turn session with a 35k-token cached prefix costs about $0.031 on Haiku 5.5 and $3.10 on GPT-6 Astra. Arithmetic.

Assumed session
An editing session on a video agent usually repeats a long prefix: system prompt, tool list, brief. Assume a 35,000-token cached prefix, 5,000 fresh input tokens per turn, 1,000 output tokens per turn, 20 turns, one cache write on turn one and cache reads on turns 2 to 20. The prompt stays at 40,000 tokens, under Haiku's 100,000-token step. These numbers are illustrative; real sessions vary.
Rates used
Haiku 5.5: input $0.10, 5-minute cache write $0.125, cache hit $0.01, output $0.50 per million (Anthropic page). Astra: input $10, cache write $12.50, cached input $1, output $50 per million (OpenAI page). Both read 2026-10-08 and match Sume's rate cards.
| Line | Haiku 5.5 | GPT-6 Astra |
|---|---|---|
| Cache write, turn 1 (35,000 tokens) | 35,000 x 0.125 / 1M = $0.004375 | 35,000 x 12.50 / 1M = $0.4375 |
| Cache reads, turns 2-20 (19 x 35,000) | 665,000 x 0.01 / 1M = $0.00665 | 665,000 x 1 / 1M = $0.665 |
| Fresh input (20 x 5,000) | 100,000 x 0.10 / 1M = $0.01 | 100,000 x 10 / 1M = $1.00 |
| Output (20 x 1,000) | 20,000 x 0.50 / 1M = $0.01 | 20,000 x 50 / 1M = $1.00 |
| Session total | $0.031025 | $3.1025 |
Reading it
Astra's total is 100 times Haiku's, the same as the ratio of their list prices, because every line scales by the same factor. What changes is where the money goes. On Astra, the one-time cache write ($0.4375) is about 14 times the whole Haiku session, and fresh input plus output are each a third of the bill. Caching saves real money only if the prefix is reused: without it, the 20 turns would send 40,000 fresh tokens each, 20 x 40,000 x 10 / 1M = $8.00 of input on Astra.
On Haiku 5.5 the 5-minute cache write is valid for five minutes (Anthropic page), so a session with long pauses pays the write again.
On Sume
Whether a Sume agent turn gets the cache discount depends on the runtime path, which the public docs do not specify. Check usage.debited_usd_micros on a finished run to see what you were actually charged. The media a session produces is billed on top and is usually the larger number.
What would change the numbers
Three things move this table. A longer prefix raises the cache read line linearly: 70,000 cached tokens double it. A prompt that grows past 100,000 tokens puts Haiku on the over-100k rates, where cache reads cost $0.05 instead of $0.01. And if the session pauses for more than five minutes between turns, the 5-minute cache entry on Haiku lapses and the write is paid again at $0.125 per million.
For Astra the page I read lists the cache write at $12.50 and cached input at $1 per million without a stated expiry, so I make no claim about how long the cache lives there.
- Doubling the prefix to 70,000: Haiku reads become 19 x 70,000 x 0.01 / 1M = $0.0133.
- Same change on Astra: 19 x 70,000 x 1 / 1M = $1.33.
Sources
Related posts
More in Models
- US company and open video weights: H3 is out, three others differ
MiniMax H3 weights exclude the US by license. HunyuanVideo, Wan 2.2 and LTX-2.5 read differently. A territory table and a hosted route on Sume.
- Veo 3.1 clip older than 2 days: it cannot be extended
Veo 3.1 extends only Veo-made clips kept 2 days. If yours expired, here is what Google's page allows and what Sume's Omni edit and new generation do.
- Video editing board: Wan 3.0 1,187, Seedance 2.5 1,151; Sume uses Omni
Wan 3.0 tops the AA video-editing board at 1,187 Elo. Sume's documented edit route is video_url on gemini-omni-flash-1.1, third at 1,139. Price inside.
- AI video with native audio: which Sume models let you turn it off
On Sume, only Kling v3 Pro prices audio separately (+$0.07/s). Gemini Omni and MiniMax always make audio; Grok Imagine and Genjutsu make none. Table inside.
Written by Sume