What do 10,000 Haiku 5.5 agent turns cost? 8k vs 120k prompts

Worked token math from Anthropic's Haiku 5.5 price page: 10,000 turns of 8k tokens cost $10.50, while 10,000 turns of 120k tokens cost $625.

5 min readSume
All posts

At Anthropic's published Claude Haiku 5.5 list prices, 10,000 agent turns with an 8,000-token prompt and 500 output tokens cost $10.50. The same 10,000 turns with a 120,000-token prompt and 1,000 output tokens cost $625.00, about 60 times more, because prompts over 100K tokens move to a higher rate.

The price page

Anthropic's Haiku page, read on 2026-10-08, lists Claude Haiku 5.5 as released on October 7, 2026. For prompts up to 100K tokens the price is $0.10 per million input tokens and $0.50 per million output tokens. For prompts over 100K tokens it is $0.50 per million input and $2.50 per million output. The page positions the model for high-volume text tasks, real-time experiences, subagents, browser automation, and simple coding.

The arithmetic

Scenario A is a lean loop: 8,000 prompt tokens and 500 output tokens per turn. Scenario B is a bloated loop that drags a long history and a large tool list into every turn: 120,000 prompt tokens and 1,000 output tokens.

Cost of 10,000 turns at Anthropic list prices, computed from the Haiku page (read 2026-10-08)
ItemScenario A (8k in, 500 out)Scenario B (120k in, 1k out)
Input tokens80 million1,200 million
Input rate$0.10 per million$0.50 per million
Input cost80 x 0.10 = $8.001,200 x 0.50 = $600.00
Output tokens5 million10 million
Output rate$0.50 per million$2.50 per million
Output cost5 x 0.50 = $2.5010 x 2.50 = $25.00
Total$10.50$625.00

What a video agent should take from it

The step at 100K is the lever. An agent that calls a hosted tool server accumulates tool results, schemas, and history. Sume's hosted MCP lets a client discover tools with tools_list and tools_schema, and the tools and gates docs tell you to use those to read the live contract. Fetch the schema of the tool you are about to call rather than loading every schema into the prompt.

The same docs recommend script_run when a turn needs three or more independent calls of the same shape. The loop runs on the Sume side and returns one value, so the repeated results do not pile into your model's prompt.

Caveats

The points that matter here, in the order you will hit them:

  • These are Anthropic's list prices and assumed token counts; they are not a quote for any Sume plan.
  • Whether the over-100K rate applies to the whole prompt or part of it is Anthropic's rule; read their page before you budget.
  • Generation spend on Sume (images, video, audio) is a separate cost, capped by generation_spend_cap_usd or max_spend_usd.

Where a loop crosses 100K

Suppose each tool result adds 6,000 tokens to the prompt and the loop starts empty. After 16 turns the prompt holds 96,000 tokens, under the line. On turn 17 it holds 102,000, over it. From there every further turn is billed at $0.50 per million input tokens instead of $0.10, a fivefold step. So the useful limit is about 16 turns of that loop, not 100K tokens in general. Trim tool output, summarize older turns, or hand repeated work to a script, before turn 17.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume