Same Sume agent run: Haiku 5.5 is $0.04, Mistral Large 4 is $0.50

One workload, two vendor price pages: 40 turns of 8,000 input and 400 output tokens driving Sume tools. Haiku 5.5: $0.04. Mistral Large 4: about $0.50.

4 min readSume
All posts

For the same assumed agent run, Claude Haiku 5.5 costs $0.04 in tokens and Mistral Large 4 costs about $0.50, a ratio of roughly 12.5 to 1. The run is 40 turns of 8,000 input and 400 output tokens each, with a client calling Sume's hosted MCP tools. Both numbers come from the vendors' own pages, read on 2026-10-08. The workload is an assumption, not a benchmark, and price is not quality.

The two price pages

Anthropic lists Claude Haiku 5.5 (API id claude-haiku-5-5, released 2026-10-07) at $0.10 input and $0.50 output per million tokens for requests up to 100k input tokens, and $0.50 input and $2.50 output above that. Mistral lists Large 4, in preview since 2026-10-06, at $1.36 input and $4.18 output per million tokens. Neither is a Sume model: Sume's docs do not list them, and they run in your own client.

Listed token prices per million tokens, read 2026-10-08
ModelInputOutputNote
Claude Haiku 5.5$0.10$0.50Up to 100k input tokens per request
Claude Haiku 5.5, larger requests$0.50$2.50Above 100k input tokens
Mistral Large 4 (preview)$1.36$4.18Preview API on Mistral Studio

Two workloads, worked out

Workload A is 40 turns of 8,000 input and 400 output tokens: 320,000 input and 16,000 output tokens in total. Workload B is 10 turns of 30,000 input and 600 output tokens: 300,000 input and 6,000 output tokens. Every request stays under 100k tokens, so Haiku uses its lower rate.

Token cost for each assumed workload at the listed rates, as of 2026-10-08
WorkloadHaiku 5.5Mistral Large 4
A: 320,000 in, 16,000 out0.32 x 0.10 + 0.016 x 0.50 = $0.0400.32 x 1.36 + 0.016 x 4.18 = $0.50208
B: 300,000 in, 6,000 out0.30 x 0.10 + 0.006 x 0.50 = $0.0330.30 x 1.36 + 0.006 x 4.18 = $0.43308

Where the gap stops mattering

Sume bills generations separately at list times 1.25, rounded up to the cent. One 5-second 720p 9:16 Seedance 2.5 clip is 289 cents. The $0.46 difference between the two models in workload A is less than a sixth of that single clip. If your agent is mostly planning and polling, the token bill is the cost that scales with the number of turns. If it mostly generates, the generation spend dominates and the cap matters more than the model.

Two Sume-side habits keep either model from overspending:

  • Pass max_spend_usd with each paid call so Sume refuses a call above your cap.
  • Use one batch jobs_wait for up to 20 job ids instead of many single waits, which cuts the number of turns, and so the tokens, in a render wave.

Caveats

The Haiku price steps up above 100k input tokens in one request, so a very long tool list or transcript changes the answer. Mistral's page does not state a tiered price or a context window, so this post does not model either. Both vendors can change rates; re-check the pages before you commit a budget.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume