Same Sume agent run: Haiku 5.5 is $0.04, Mistral Large 4 is $0.50
One workload, two vendor price pages: 40 turns of 8,000 input and 400 output tokens driving Sume tools. Haiku 5.5: $0.04. Mistral Large 4: about $0.50.

For the same assumed agent run, Claude Haiku 5.5 costs $0.04 in tokens and Mistral Large 4 costs about $0.50, a ratio of roughly 12.5 to 1. The run is 40 turns of 8,000 input and 400 output tokens each, with a client calling Sume's hosted MCP tools. Both numbers come from the vendors' own pages, read on 2026-10-08. The workload is an assumption, not a benchmark, and price is not quality.
The two price pages
Anthropic lists Claude Haiku 5.5 (API id claude-haiku-5-5, released 2026-10-07) at $0.10 input and $0.50 output per million tokens for requests up to 100k input tokens, and $0.50 input and $2.50 output above that. Mistral lists Large 4, in preview since 2026-10-06, at $1.36 input and $4.18 output per million tokens. Neither is a Sume model: Sume's docs do not list them, and they run in your own client.
| Model | Input | Output | Note |
|---|---|---|---|
| Claude Haiku 5.5 | $0.10 | $0.50 | Up to 100k input tokens per request |
| Claude Haiku 5.5, larger requests | $0.50 | $2.50 | Above 100k input tokens |
| Mistral Large 4 (preview) | $1.36 | $4.18 | Preview API on Mistral Studio |
Two workloads, worked out
Workload A is 40 turns of 8,000 input and 400 output tokens: 320,000 input and 16,000 output tokens in total. Workload B is 10 turns of 30,000 input and 600 output tokens: 300,000 input and 6,000 output tokens. Every request stays under 100k tokens, so Haiku uses its lower rate.
| Workload | Haiku 5.5 | Mistral Large 4 |
|---|---|---|
| A: 320,000 in, 16,000 out | 0.32 x 0.10 + 0.016 x 0.50 = $0.040 | 0.32 x 1.36 + 0.016 x 4.18 = $0.50208 |
| B: 300,000 in, 6,000 out | 0.30 x 0.10 + 0.006 x 0.50 = $0.033 | 0.30 x 1.36 + 0.006 x 4.18 = $0.43308 |
Where the gap stops mattering
Sume bills generations separately at list times 1.25, rounded up to the cent. One 5-second 720p 9:16 Seedance 2.5 clip is 289 cents. The $0.46 difference between the two models in workload A is less than a sixth of that single clip. If your agent is mostly planning and polling, the token bill is the cost that scales with the number of turns. If it mostly generates, the generation spend dominates and the cap matters more than the model.
Two Sume-side habits keep either model from overspending:
- Pass
max_spend_usdwith each paid call so Sume refuses a call above your cap. - Use one batch
jobs_waitfor up to 20 job ids instead of many single waits, which cuts the number of turns, and so the tokens, in a render wave.
Caveats
The Haiku price steps up above 100k input tokens in one request, so a very long tool list or transcript changes the answer. Mistral's page does not state a tiered price or a context window, so this post does not model either. Both vendors can change rates; re-check the pages before you commit a budget.
Sources
Related posts
More in Agents
- Waiting 10 minutes for a Sume render over MCP: 11 jobs_wait calls
A remote MCP jobs_wait holds at most 55 seconds, so a 10-minute render takes up to 11 calls at the cap, 12 at the default 50. Re-issue the wait; never resubmit.
- Mistral Large 4 as your agent: 40 Sume tool turns cost $0.50
Mistral lists Large 4 at $1.36 in and $4.18 out per million tokens. For a 40-turn agent that drives Sume, the token bill is $0.50, next to a $3.47 clip.
- Nightly Sume schedule: model choice and the 30-day cap math
A nightly Scheduled run caps generation at $1.00 by default, so 30 nights is at most $30. How the model's token cost sits beside that cap, with the arithmetic.
- One Sume webhook endpoint for three run types: verify, route
Format, action and agent runs share one signature scheme. Verify HMAC-SHA256 over timestamp.body, refuse an empty secret, then route on the event field.
Written by Sume