A 20-turn agent session: Haiku 5.5 $0.03, Astra $3.10 with caching

Cache write on turn one, reads after: a 20-turn session with a 35k-token cached prefix costs about $0.031 on Haiku 5.5 and $3.10 on GPT-6 Astra. Arithmetic.

4 min readSume
All posts

Assumed session

An editing session on a video agent usually repeats a long prefix: system prompt, tool list, brief. Assume a 35,000-token cached prefix, 5,000 fresh input tokens per turn, 1,000 output tokens per turn, 20 turns, one cache write on turn one and cache reads on turns 2 to 20. The prompt stays at 40,000 tokens, under Haiku's 100,000-token step. These numbers are illustrative; real sessions vary.

Rates used

Haiku 5.5: input $0.10, 5-minute cache write $0.125, cache hit $0.01, output $0.50 per million (Anthropic page). Astra: input $10, cache write $12.50, cached input $1, output $50 per million (OpenAI page). Both read 2026-10-08 and match Sume's rate cards.

20-turn session arithmetic (rates read 2026-10-08)
LineHaiku 5.5GPT-6 Astra
Cache write, turn 1 (35,000 tokens)35,000 x 0.125 / 1M = $0.00437535,000 x 12.50 / 1M = $0.4375
Cache reads, turns 2-20 (19 x 35,000)665,000 x 0.01 / 1M = $0.00665665,000 x 1 / 1M = $0.665
Fresh input (20 x 5,000)100,000 x 0.10 / 1M = $0.01100,000 x 10 / 1M = $1.00
Output (20 x 1,000)20,000 x 0.50 / 1M = $0.0120,000 x 50 / 1M = $1.00
Session total$0.031025$3.1025

Reading it

Astra's total is 100 times Haiku's, the same as the ratio of their list prices, because every line scales by the same factor. What changes is where the money goes. On Astra, the one-time cache write ($0.4375) is about 14 times the whole Haiku session, and fresh input plus output are each a third of the bill. Caching saves real money only if the prefix is reused: without it, the 20 turns would send 40,000 fresh tokens each, 20 x 40,000 x 10 / 1M = $8.00 of input on Astra.

On Haiku 5.5 the 5-minute cache write is valid for five minutes (Anthropic page), so a session with long pauses pays the write again.

On Sume

Whether a Sume agent turn gets the cache discount depends on the runtime path, which the public docs do not specify. Check usage.debited_usd_micros on a finished run to see what you were actually charged. The media a session produces is billed on top and is usually the larger number.

What would change the numbers

Three things move this table. A longer prefix raises the cache read line linearly: 70,000 cached tokens double it. A prompt that grows past 100,000 tokens puts Haiku on the over-100k rates, where cache reads cost $0.05 instead of $0.01. And if the session pauses for more than five minutes between turns, the 5-minute cache entry on Haiku lapses and the write is paid again at $0.125 per million.

For Astra the page I read lists the cache write at $12.50 and cached input at $1 per million without a stated expiry, so I make no claim about how long the cache lives there.

  • Doubling the prefix to 70,000: Haiku reads become 19 x 70,000 x 0.01 / 1M = $0.0133.
  • Same change on Astra: 19 x 70,000 x 1 / 1M = $1.33.

Sources

Related posts

More in Models

All Models posts

Written by Sume