Mistral Large 4 on a 12-turn storyboard: tokens are 4% of the bill

Mistral Large 4 lists $1.36 input and $4.18 output per million tokens. A 12-turn run that makes six Wan 3.0 clips spends 4.3% of its total on tokens.

5 min readSume
All posts

If Mistral Large 4 drives a Sume storyboard, the model tokens are a small slice of the bill and the generated clips are most of it. In the example below, 12 turns cost $0.166944 in Mistral tokens, while six five-second Wan 3.0 720p clips cost $3.75, so tokens are 4.26% of the $3.916944 total. Swapping to a cheaper model moves the total by cents; capping generation moves it by dollars.

Mistral's launch page, read 2026-10-09, lists Large 4 at $1.36 per million input tokens and $4.18 per million output tokens through a public preview API, with open weights due by the end of October 2026. It describes a 1-trillion-parameter multimodal model with 52 billion active parameters. The page does not describe MCP support, so whether a given Mistral client can connect to a remote MCP server is a question for that client. Sume's hosted endpoint is a standard remote MCP server at https://mcp.sume.com/mcp.

Assumptions and arithmetic

The scenario is invented for the arithmetic, not measured: 12 turns, 9,000 input tokens and 400 output tokens per turn, with no history growth. Replace those with your own logs. The Haiku 5.5 column uses Anthropic's lower-tier rates ($0.10 and $0.50 per million) for comparison.

Token cost vs generation cost for one invented 12-turn run (vendor prices read 2026-10-09; Wan 3.0 720p from the Sume catalog)
LineMistral Large 4Claude Haiku 5.5
Input, 12 x 9,000 = 108,000 tokens108,000 x $1.36 / 1M = $0.14688108,000 x $0.10 / 1M = $0.0108
Output, 12 x 400 = 4,800 tokens4,800 x $4.18 / 1M = $0.0200644,800 x $0.50 / 1M = $0.0024
Token total$0.166944$0.0132
Six Wan 3.0 720p clips, 5 s each6 x 5 x $0.125 = $3.756 x 5 x $0.125 = $3.75
Run total$3.916944$3.7632
Token share of total4.26%0.35%

What to cap

The model choice changes $0.15 here. The clip count changes $0.625 per clip. For an unattended agent the sensible controls sit on the Sume side:

  • Call dry_run=true or generation_admission_preview before the first paid submit; see Tools and gates.
  • Pass max_spend_usd on paid calls. Sume enforces it only when you provide it, so a prompt that forgets the field has no cap from this field.
  • Send a fresh idempotency_key per intended clip and reuse it only for an exact retry.
  • If the agent runs on Sume's side instead, Agent Completions require generation_spend_cap_usd on every request.

Where this comparison stops

Sume does not list Mistral Large 4 as a selectable agent model; Agent Completions accept only sume-agent. The comparison is for people who run their own Mistral-driven client against the hosted MCP endpoint. Token counts, tier changes above 100,000 tokens, and any caching discount should be checked against the vendor page before you budget.

Reading the table

The useful number is the last row. Even on the pricier model, tokens stay under five percent of this run. Doubling the turn count to 24 would add another $0.166944 on Mistral, still far below one extra clip at $0.625. The share only flips when an agent loops without making anything, for example retrying a refused call dozens of times, which is exactly the case an idempotency key and a spend cap are for.

If you want a number for your own run, count tokens per turn from your client's usage log, multiply by the two vendor prices, and add the Sume generation line from the public catalog. Keep the arithmetic visible in your runbook so a price change on either side is easy to re-run.

Sources

Related posts

More in Models

All Models posts

Written by Sume