Which video model does a Sume Format use? Not the model field

The model field on a Sume Format run picks the orchestrating LLM, default gpt-6-sol. The Format's tools choose the image, video and audio models.

4 min readSume
All posts

The model field on a Sume Format run does not choose the video model. According to the docs, it selects only the orchestrator LLM, the one that reads the Format and drives the tools, and the default is gpt-6-sol. The image, video and audio models come from the tools the Format calls.

Who picks what

Mixing these up is the most common reason a run does not look the way you expected. Passing a video model id as model will not change the clip.

Who selects what in a Format run, from the Sume docs, read 2026-10-06.
LayerChosen byHow to influence it
Orchestrator LLMmodel on the run, default gpt-6-solSet model
Video modelThe Format's tool callsName your needs in instruction
Image and audio modelsThe Format's tool callsSame
A specific video modelYou, outside FormatsCall POST /v1/videos directly

If you need a specific model

Go to the model endpoints. GET /v1/videos/models lists each model with its supported resolutions, aspect ratios, durations and pricing SKUs, and POST /v1/videos takes the model you name. That is the right path when a brief depends on one model's strength, such as a 30-second pass or a particular ratio.

A Format is a recipe, and the recipe decides its tools. You can steer with instruction, which is placed after the Format body so it wins where they disagree, but the docs do not promise that a request for a named model will be honored.

Read the receipt to see what ran

After a run, usage shows what it cost. usage.debited_usd_micros is the real cost, while billable_amount_usd_micros leaves out the orchestrator turn. If the model matters to you, compare runs and their costs before you standardize on a Format.

Summary

Short version: model is the planner, the Format is the recipe, and the video endpoints are where you choose a video model yourself.

A common mistake

People often send a model name they saw in a blog post in the model field and expect different video. The orchestrator is a language model, so a video model id in that field is not what you meant. If your brief is model-specific, use the video endpoints, and keep Formats for the jobs where the recipe is the point.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume