1,000 agent turns: Haiku 5.5 about $4, GPT-6 Astra about $400

At list rates, 1,000 turns of 30,000 input and 2,000 output tokens cost about $4 on Haiku 5.5 and $400 on GPT-6 Astra. The arithmetic and the cache effect.

4 min readSume
All posts

The bill

For a thousand turns that each send 30,000 fresh input tokens and return 2,000 output tokens, Claude Haiku 5.5 costs about $4 and GPT-6 Astra costs about $400. That is a 100x gap, and it comes straight from the list rates on the Anthropic and OpenAI pages read on 2026-10-08. The turn shape is an assumption for illustration; Sume's own rate cards use the same prices.

Sume's registry also lists GLM 5.3 Flash. Its Sume card is lower than the Z.ai page shows, so it is shown separately below with both numbers.

Table

Per-turn cost is input tokens times the input rate plus output tokens times the output rate, divided by one million.

1,000 turns of 30,000 input and 2,000 output tokens, no cache (rates read 2026-10-08)
ModelInput / output $ per MPer turnx 1,000
Haiku 5.5 (Anthropic page)0.10 / 0.500.0030 + 0.0010 = $0.0040$4.00
GLM-5.3-Flash (Z.ai page)0.15 / 0.500.0045 + 0.0010 = $0.0055$5.50
GLM 5.3 Flash (Sume card in registry)0.075 / 0.250.00225 + 0.0005 = $0.00275$2.75
GPT-6 Astra (OpenAI page)10 / 500.30 + 0.10 = $0.40$400.00
GPT-6 Astra Fast (2x in Sume)20 / 1000.60 + 0.20 = $0.80$800.00

Caching changes the picture, not the order

If 25,000 of the 30,000 input tokens are cache reads, Haiku's input drops to 25,000 x 0.01 / 1M = $0.00025 plus 5,000 x 0.10 / 1M = $0.0005, so the turn is $0.00025 + $0.0005 + $0.0010 = $0.00175. Astra's cached input is $1 per million: 25,000 x 1 / 1M = $0.025 plus 5,000 x 10 / 1M = $0.05 plus $0.10 output gives $0.175. The ratio stays near 100x.

Output is the stubborn part. On Astra the 2,000 output tokens alone are $0.10, a quarter of the uncached turn.

What it does not include

These figures are model tokens only. The media a video agent generates is billed separately and is the larger line on most runs: a single 10 second Seedance 2 clip at 720p is already $3.78 at Sume's price. The agent's model choice moves the smaller line, and moves it by cents or by tens of cents.

Where the gap narrows

A 100x gap is the headline, but it is not fixed. It shrinks when Haiku's turns cross the 100,000-token step. A turn of 150,000 input and 3,000 output tokens costs 150,000 x 0.50 / 1M = $0.075 plus 3,000 x 2.50 / 1M = $0.0075, so $0.0825 on Haiku, against 150,000 x 10 / 1M = $1.50 plus 3,000 x 50 / 1M = $0.15, so $1.65 on Astra. That is a 20x gap. So a lean Haiku prompt is worth keeping under 100,000 tokens.

The gap also shrinks if Astra's work removes retries. If Haiku needs three attempts where Astra needs one, the $4 per thousand becomes $12, which is still 33x cheaper. I make no claim about how often that happens; Anthropic positions Haiku 5.5 for classification, extraction and routing, and I have not benchmarked either model.

  • Under 100k prompt: 100x gap at the assumed shape.
  • At 150k prompt: about 20x.

Sources

Related posts

More in Models

All Models posts

Written by Sume