Mistral Large 4 as your agent: 40 Sume tool turns cost $0.50

Mistral lists Large 4 at $1.36 in and $4.18 out per million tokens. For a 40-turn agent that drives Sume, the token bill is $0.50, next to a $3.47 clip.

4 min readSume
All posts

If you let Mistral Large 4 drive an agent that calls Sume's hosted MCP tools, the model tokens for a 40-turn session come to about $0.50 under the workload assumed below. That is $1.36 per million input tokens and $4.18 per million output tokens, the rates on Mistral's announcement page. The Sume side is billed separately, per generation, and one 6-second vertical clip costs more than the whole token bill.

Mistral Large 4 is a vendor model that you run in your own client. Sume's docs do not list it as a model, so Sume does not route to it or bill for it. Sume only sees the MCP tool calls the client sends. The assumptions in this post are mine, not a benchmark: change them to match your own logs.

What Mistral publishes (read 2026-10-08)

Mistral announced Large 4 on 2026-10-06 as a public preview. The page says it is a 1 trillion-parameter natively multimodal model with 52 billion active parameters, available through a preview API on Mistral Studio, with weights planned for the end of the month. The page does not state a context window, so this post makes no context claim.

Mistral Large 4 preview facts, from the vendor page, read 2026-10-08
ItemValue
Preview date2026-10-06
Input price$1.36 per million tokens
Output price$4.18 per million tokens
Size1 trillion parameters, 52 billion active
WeightsPlanned for the end of the month

The arithmetic for 40 turns

Assume the agent runs 40 turns. Each turn sends 8,000 input tokens (the conversation so far, plus the tool list) and returns 400 output tokens (a tool call or a short reply). Those figures are an assumption for a mid-size session, not a measurement.

Token bill for the assumed 40-turn session at Mistral's listed rates, as of 2026-10-08
LineTokensRate per millionCost
Input40 x 8,000 = 320,000$1.360.32 x 1.36 = $0.4352
Output40 x 400 = 16,000$4.180.016 x 4.18 = $0.06688
Total336,000$0.50208, about $0.50

Compare it with one Sume clip

Sume bills a generation at the catalog list price times 1.25, rounded up to the cent. A 6-second 720p 9:16 clip on Seedance 2.5 is 347 cents, or $3.47. The 40-turn token bill is about 14.5 percent of that one clip ($0.50208 divided by $3.47). If the session ends with seven such clips, the token bill is about 2 percent of the $24.29 generation spend.

So for a generation-heavy agent, the model choice matters less than the spend controls. Use the documented gates on every paid call:

  • Send a fresh idempotency_key on each paid call.
  • Run dry_run first to see the estimate before anything is submitted.
  • Set max_spend_usd when you want Sume to enforce a cap; Sume enforces it only when you provide it.

What to check before you switch models

The preview label matters: Mistral says the preview is being red-teamed with partners before a wider release, so rates and availability can change. Re-read the page before you budget. Also count your own turns. A long tool list in every request raises the 8,000-token figure, and the 40-turn total scales linearly with it.

Connecting is the same for any client: point it at https://mcp.sume.com/mcp, use OAuth for interactive work or an API key for automation, and call tools_list once to confirm the session.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume