1,000 agent turns: Haiku 5.5 about $4, GPT-6 Astra about $400
At list rates, 1,000 turns of 30,000 input and 2,000 output tokens cost about $4 on Haiku 5.5 and $400 on GPT-6 Astra. The arithmetic and the cache effect.

The bill
For a thousand turns that each send 30,000 fresh input tokens and return 2,000 output tokens, Claude Haiku 5.5 costs about $4 and GPT-6 Astra costs about $400. That is a 100x gap, and it comes straight from the list rates on the Anthropic and OpenAI pages read on 2026-10-08. The turn shape is an assumption for illustration; Sume's own rate cards use the same prices.
Sume's registry also lists GLM 5.3 Flash. Its Sume card is lower than the Z.ai page shows, so it is shown separately below with both numbers.
Table
Per-turn cost is input tokens times the input rate plus output tokens times the output rate, divided by one million.
| Model | Input / output $ per M | Per turn | x 1,000 |
|---|---|---|---|
| Haiku 5.5 (Anthropic page) | 0.10 / 0.50 | 0.0030 + 0.0010 = $0.0040 | $4.00 |
| GLM-5.3-Flash (Z.ai page) | 0.15 / 0.50 | 0.0045 + 0.0010 = $0.0055 | $5.50 |
| GLM 5.3 Flash (Sume card in registry) | 0.075 / 0.25 | 0.00225 + 0.0005 = $0.00275 | $2.75 |
| GPT-6 Astra (OpenAI page) | 10 / 50 | 0.30 + 0.10 = $0.40 | $400.00 |
| GPT-6 Astra Fast (2x in Sume) | 20 / 100 | 0.60 + 0.20 = $0.80 | $800.00 |
Caching changes the picture, not the order
If 25,000 of the 30,000 input tokens are cache reads, Haiku's input drops to 25,000 x 0.01 / 1M = $0.00025 plus 5,000 x 0.10 / 1M = $0.0005, so the turn is $0.00025 + $0.0005 + $0.0010 = $0.00175. Astra's cached input is $1 per million: 25,000 x 1 / 1M = $0.025 plus 5,000 x 10 / 1M = $0.05 plus $0.10 output gives $0.175. The ratio stays near 100x.
Output is the stubborn part. On Astra the 2,000 output tokens alone are $0.10, a quarter of the uncached turn.
What it does not include
These figures are model tokens only. The media a video agent generates is billed separately and is the larger line on most runs: a single 10 second Seedance 2 clip at 720p is already $3.78 at Sume's price. The agent's model choice moves the smaller line, and moves it by cents or by tens of cents.
Where the gap narrows
A 100x gap is the headline, but it is not fixed. It shrinks when Haiku's turns cross the 100,000-token step. A turn of 150,000 input and 3,000 output tokens costs 150,000 x 0.50 / 1M = $0.075 plus 3,000 x 2.50 / 1M = $0.0075, so $0.0825 on Haiku, against 150,000 x 10 / 1M = $1.50 plus 3,000 x 50 / 1M = $0.15, so $1.65 on Astra. That is a 20x gap. So a lean Haiku prompt is worth keeping under 100,000 tokens.
The gap also shrinks if Astra's work removes retries. If Haiku needs three attempts where Astra needs one, the $4 per thousand becomes $12, which is still 33x cheaper. I make no claim about how often that happens; Anthropic positions Haiku 5.5 for classification, extraction and routing, and I have not benchmarked either model.
- Under 100k prompt: 100x gap at the assumed shape.
- At 150k prompt: about 20x.
Sources
Related posts
More in Models
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
- Higgsfield Soul on Sume: 1 cent an image, n of 1 or 4, $100 per 10,000
Higgsfield Soul is the cheapest Sume image model at 1 cent. It is text-only, takes n of 1 or 4 and 720p or 1080p, so 10,000 images is 2,500 calls and $100.
- High to max on 100 hero images: GPT Image 2.5 adds $20, xhigh adds $5
Upgrade 100 GPT Image 2.5 hero images from high to xhigh or max: the extra cost on Sume is $5 or $20 at 1024-class size. The arithmetic is shown.
- How long can a Sume agent run take? Timeouts for slow models
Format runs are finalized at 90 minutes, or after 10 silent minutes once past 25. Agent Completions publish no deadline. Webhooks time out at 10 s per attempt.
Written by Sume