Which video model does a Sume Format use? Not the model field
The model field on a Sume Format run picks the orchestrating LLM, default gpt-6-sol. The Format's tools choose the image, video and audio models.

The model field on a Sume Format run does not choose the video model. According to the docs, it selects only the orchestrator LLM, the one that reads the Format and drives the tools, and the default is gpt-6-sol. The image, video and audio models come from the tools the Format calls.
Who picks what
Mixing these up is the most common reason a run does not look the way you expected. Passing a video model id as model will not change the clip.
| Layer | Chosen by | How to influence it |
|---|---|---|
| Orchestrator LLM | model on the run, default gpt-6-sol | Set model |
| Video model | The Format's tool calls | Name your needs in instruction |
| Image and audio models | The Format's tool calls | Same |
| A specific video model | You, outside Formats | Call POST /v1/videos directly |
If you need a specific model
Go to the model endpoints. GET /v1/videos/models lists each model with its supported resolutions, aspect ratios, durations and pricing SKUs, and POST /v1/videos takes the model you name. That is the right path when a brief depends on one model's strength, such as a 30-second pass or a particular ratio.
A Format is a recipe, and the recipe decides its tools. You can steer with instruction, which is placed after the Format body so it wins where they disagree, but the docs do not promise that a request for a named model will be honored.
Read the receipt to see what ran
After a run, usage shows what it cost. usage.debited_usd_micros is the real cost, while billable_amount_usd_micros leaves out the orchestrator turn. If the model matters to you, compare runs and their costs before you standardize on a Format.
Summary
Short version: model is the planner, the Format is the recipe, and the video endpoints are where you choose a video model yourself.
A common mistake
People often send a model name they saw in a blog post in the model field and expect different video. The orchestrator is a language model, so a video model id in that field is not what you meant. If your brief is model-specific, use the video endpoints, and keep Formats for the jobs where the recipe is the point.
Sources
Related posts
More in Formats
- Why did my Format run do that? Read the first message of its thread
A Format run's first thread message is the text the agent received: Format pointer, your instruction, unattended note and input file path. Read it first.
- Ready-made Formats for product video: the Sume Format catalog
Sume ships ready-made Formats for product and UGC-style video and images, each callable from your backend with one HTTP request at the reserved sume handle.
- What is a Sume Format? Turn an agent thread into one API call
A Sume Format is a saved video recipe your backend calls by handle and slug. One POST runs it in a fresh sandbox and returns media plus optional typed JSON.
- How to embed AI video generation in your product with Sume Formats
To embed AI video generation, your server holds one Sume API key and runs a Format per customer, with a derived Idempotency-Key, spend cap, and webhook.
Written by Sume