Which LLM should drive an AI video agent on Sume? A catalog table
Five orchestrator families sit in Sume's Agents catalog. The vendor-page limits that matter for a video agent, and what the LLM does and does not decide.

For an AI video agent on Sume, start with the default orchestrator and change it only when a run shows a reason. The LLM plans and calls tools; the video, image and audio models are chosen by the Format's tools, so swapping GPT-6.1 Sol for Sonnet 5.5 does not change which model renders a clip.
What does the LLM decide in a video agent?
It reads your instruction and input, looks at attached images, decides which tools to call and with what arguments, and writes the structured output. It does not render. The Create a run page says the model field "selects the orchestrator only".
That sets what to compare: how reliably it follows a brief, how it handles up to 30 attached images, how long its plan takes, and what its own tokens cost. Picking by general benchmark rank tells you little about those, and this post does not quote one.
How do the candidates compare on vendor pages?
Only limits and list prices the vendors state themselves. Sume's rate for a run is on the receipt:
- GPT-6.1 Sol's page does not state image input in the fields I read, so the table does not either.
- Not every row is listed everywhere: the Anthropic rows other than Opus 5.5, DeepSeek and Grok 4.7 depend on catalog gates in your environment.
| Sume row | Vendor context window | Image input | List price per 1M tokens, input / output |
|---|---|---|---|
| GPT-6.1 Sol | 1,050,000 | Not stated on the page I read | $2 / $10 |
| Claude Opus 5.5 | 1M | Text and image | $4 / $20 |
| Claude Sonnet 5.5 | 1M | Text and image | $2 / $10 |
| Grok 4.7 | 500,000 | Text and image | $2 / $6 under 200k tokens |
| DeepSeek V4.1 Flash | 1M | Vision supported | $0.6 to $1.2 output; input $0.15 to $0.3 uncached |
Which one should I try first?
Leave model off and read the receipt. On current main a Format run with no model uses GPT-6.1 Sol, per the default post, so you start from a known orchestrator without choosing one.
Then change one variable at a time. If plans miss details, try Opus 5.5 or raise effort. If the orchestrator's cost matters more than polish, try Sonnet 5.5 or DeepSeek V4.1 Flash. If you need Grok for another reason, check the Grok alias post first so an old id does not give you a GPT.
What about long briefs and many images?
Check the limits that sit in front of the model before the model itself. A Format run accepts an instruction of up to 8000 characters, of which about the first 4000 are carried to the run as prompt text, per the Create a run page, so a long brief belongs in input, which reaches the run in full as a file. Up to 30 images can be attached, and media URLs inside input share that budget.
Those limits are the same whichever orchestrator you pick. A 1M window does not let you send more than the API accepts, so choosing a larger model to fit a longer prompt solves the wrong problem.
How do I compare two orchestrators fairly?
Same Format, same input, same output_schema, a separate Idempotency-Key for each, and a low generation_spend_cap_usd so a bad plan is cheap. Log the receipt's model, usage and the time to a terminal status.
Judge the output, not the plan: did the structured result validate, did the clip match the brief, how many media jobs did it start. One run per model is an anecdote; ten is a sample. Sume does not publish a ranking of orchestrators for video, and the table above only tells you the limits each vendor states.
Sources
Related posts
More in Use cases
- Wix Stores wants 3000x3000 images: check your model's cap
Wix Stores lists 3000 x 3000 px as its minimum product image. Check GET /v1/images/models for each model's largest size before you generate for it.
- X Ads card video: 10 minutes, 500 MB, MP4 or MOV, 16:9 or 1:1
X's Ads API docs list card video up to 10 minutes and 500 MB, MP4 or MOV, 16:9 or 1:1, and images up to 3 MB. Here is how to fit a Sume clip to that.
- YouTube ads need a business name and logo from Oct 12, 2026
From October 12, 2026 YouTube and partners line items in DV360 need a business name (25 characters) and a square logo. Specs, and why not to generate the logo.
- What happens if I don't label AI video on YouTube? Penalties
YouTube says missing an AI disclosure can lead to a manual label, removal or Partner Program suspension. What triggers it and how Sume users avoid it.
Written by Sume