Mistral Large 4 preview as your orchestrator, Sume as the renderer

Sume does not list Mistral Large 4. Let it plan in your own code and hand each video task to a Sume Agent Completion; Python sample with the required cap.

4 min readSume
All posts

Short answer

Sume does not list Mistral Large 4 in its agent model registry, and an Agent Completion accepts only model: "sume-agent". If you want Mistral's model to plan a video job, run that model in your own code and let it call Sume's HTTP API for the work. The Mistral model page, read 2026-10-08, shows the API id mistral-large-4, the status Public Preview (v26.10), a 1M-token context and 52B active of 1.05T total parameters.

I am leaving Mistral's price out. The page figures did not read consistently when fetched, and a price in a blog post should come from a page you can check today.

The split

The planner is yours. The agent that actually makes media is Sume's. Your planner produces one self-contained instruction per shot or per asset, and your code posts it.

Who does what in this setup (Sume docs and Mistral page, read 2026-10-08)
StepRuns onSource
Write the shot listMistral Large 4 preview, called from your codeMistral model page: id mistral-large-4
Run one task with tools and media generationSume Agent, POST /v1/agent/completionsAgent Completions docs
Hard limit on spendgeneration_spend_cap_usd, required on every callAgent Completions docs
ResultPoll status_url until next_action is not poll_statusAgent Completions docs

Starting one task

The create call returns 202 and a receipt. The spend cap has no default, so omit it and you get 400 invalid_request. This sample sends one instruction and prints the polling URL.

import json, os, urllib.request

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("SUME_API_KEY is empty")

body = {
    "instruction": "Make a 6 second vertical product teaser from the attached brief.",
    "generation_spend_cap_usd": 3,
}
req = urllib.request.Request(
    "https://api.sume.com/v1/agent/completions",
    data=json.dumps(body).encode(),
    headers={
        "Authorization": f"Bearer {key}",
        "Content-Type": "application/json",
        "Idempotency-Key": "teaser-001",
    },
    method="POST",
)
with urllib.request.urlopen(req) as resp:
    receipt = json.load(resp)["data"]
print(receipt["id"], receipt["status_url"])

Limits to plan around

Agent Completions are async: no streaming and no choices[] response. Assistant turns in messages[] are rejected, and each completion starts a new thread, so your planner must put all context into each instruction. Attachments are limited to 30 images.

Why keep the planner outside

There are two reasons to run the planner yourself instead of waiting for a listing. You control the model id and can pin it in your own config, which matters for a model whose Mistral page status is Public Preview (v26.10). And you control retries: if the preview changes between runs, your code sees it.

The cost of the split is that Sume's agent starts each completion in a fresh thread, so continuity between shots has to come from your planner's output. Put the shot list, the style notes and the references into the instruction or the input object every time. Sume writes input to a workspace file and treats it as data only.

A last practical point is about failure. If your own planner call fails or returns something malformed, do not start the Sume run at all. Validate the planner output against a small schema first, and only then create the completion with the cap and the idempotency key. That keeps a bad plan from turning into media spend, and it means every Sume run you pay for began from an instruction you already checked. The cap is the safety net; the validation is the first line of defense.

  • Always send generation_spend_cap_usd; a missing cap is a 400.
  • Send an Idempotency-Key per task so a retry does not start a second run.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume