Haiku 5.5 vs Sonnet 5.5 token prices for an agent calling Sume

Haiku 5.5 costs $0.10 per million input tokens and Sonnet 5.5 costs $2. What the gap does and does not change for a Sume agent's spend cap.

5 min readSume
All posts

On Anthropic's pricing page read on 2026-10-08, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and Claude Sonnet 5.5 costs $2 and $10. That is a 20 times gap on both input and output. For an agent that drives Sume's hosted MCP or Agent Completions, the gap changes the model bill for planning and polling. It does not change what Sume meters for media generation, which is capped separately.

The two models side by side

Above 100,000 tokens of prompt, Haiku 5.5 moves to $0.50 input and $2.50 output, which is the row to watch in a long polling session.

Claude API price per million tokens, Anthropic pricing page, read 2026-10-08
ItemHaiku 5.5 (prompt up to 100k)Sonnet 5.5
Base input$0.10$2
Cache hit$0.01$0.10
Output$0.50$10
Batch input / output$0.05 / $0.25$1 / $5
Tool-use system prompt (auto)286 tokens286 tokens
Default effortmediumhigh

What Sume meters

Sume's Agent Completions run the Sume agent in Sume's runtime, so the model tokens there are not on your Anthropic bill. The spend you control on that route is generation_spend_cap_usd, a required field with no default, and the model field accepts only sume-agent. On the hosted MCP route, your own model client calls Sume tools, and max_spend_usd is the optional per-call ceiling.

So the choice between Haiku and Sonnet applies when you run your own agent against hosted MCP. A cheaper planner means a cheaper loop around the same job. It does not make a Sume render cheaper.

When the cheaper model is the wrong call

Anthropic positions Haiku 5.5 for high-volume, latency-sensitive work and suggests comparing xhigh or max on Haiku with Sonnet 5.5 on performance, cost and speed. A paid-call loop with strict rules about idempotency and caps is the sort of work where a skipped check costs more than the token savings. Run your own evals on three counts: calls missing max_spend_usd, calls missing dry_run before a first submit, and duplicate submits.

If both models pass, use the cheaper one. If one fails, raise effort before switching model, since effort is the lever Anthropic documents for that. Cap sizing on the Sume side is covered in spend cap sizing for an agent completion.

What to put in the budget

Budget three lines. The first is model tokens on your Anthropic bill. The second is Sume generation spend, bounded per call by max_spend_usd or per run by generation_spend_cap_usd. The third is any runtime charge if you host the loop on a metered agent service.

The first line is the one that moves when you swap Haiku for Sonnet. The second does not. If the model choice changes how many Sume calls the agent makes, for example because a weaker planner retries a render, the second line moves too, and that effect can outweigh the token savings.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume