Which LLM should drive an AI video agent on Sume? A catalog table

Five orchestrator families sit in Sume's Agents catalog. The vendor-page limits that matter for a video agent, and what the LLM does and does not decide.

6 min readSume
All posts

For an AI video agent on Sume, start with the default orchestrator and change it only when a run shows a reason. The LLM plans and calls tools; the video, image and audio models are chosen by the Format's tools, so swapping GPT-6.1 Sol for Sonnet 5.5 does not change which model renders a clip.

What does the LLM decide in a video agent?

It reads your instruction and input, looks at attached images, decides which tools to call and with what arguments, and writes the structured output. It does not render. The Create a run page says the model field "selects the orchestrator only".

That sets what to compare: how reliably it follows a brief, how it handles up to 30 attached images, how long its plan takes, and what its own tokens cost. Picking by general benchmark rank tells you little about those, and this post does not quote one.

How do the candidates compare on vendor pages?

Only limits and list prices the vendors state themselves. Sume's rate for a run is on the receipt:

  • GPT-6.1 Sol's page does not state image input in the fields I read, so the table does not either.
  • Not every row is listed everywhere: the Anthropic rows other than Opus 5.5, DeepSeek and Grok 4.7 depend on catalog gates in your environment.
Vendor model pages (read 2026-10-02)
Sume rowVendor context windowImage inputList price per 1M tokens, input / output
GPT-6.1 Sol1,050,000Not stated on the page I read$2 / $10
Claude Opus 5.51MText and image$4 / $20
Claude Sonnet 5.51MText and image$2 / $10
Grok 4.7500,000Text and image$2 / $6 under 200k tokens
DeepSeek V4.1 Flash1MVision supported$0.6 to $1.2 output; input $0.15 to $0.3 uncached

Which one should I try first?

Leave model off and read the receipt. On current main a Format run with no model uses GPT-6.1 Sol, per the default post, so you start from a known orchestrator without choosing one.

Then change one variable at a time. If plans miss details, try Opus 5.5 or raise effort. If the orchestrator's cost matters more than polish, try Sonnet 5.5 or DeepSeek V4.1 Flash. If you need Grok for another reason, check the Grok alias post first so an old id does not give you a GPT.

What about long briefs and many images?

Check the limits that sit in front of the model before the model itself. A Format run accepts an instruction of up to 8000 characters, of which about the first 4000 are carried to the run as prompt text, per the Create a run page, so a long brief belongs in input, which reaches the run in full as a file. Up to 30 images can be attached, and media URLs inside input share that budget.

Those limits are the same whichever orchestrator you pick. A 1M window does not let you send more than the API accepts, so choosing a larger model to fit a longer prompt solves the wrong problem.

How do I compare two orchestrators fairly?

Same Format, same input, same output_schema, a separate Idempotency-Key for each, and a low generation_spend_cap_usd so a bad plan is cheap. Log the receipt's model, usage and the time to a terminal status.

Judge the output, not the plan: did the structured result validate, did the clip match the brief, how many media jobs did it start. One run per model is an anecdote; ten is a sample. Sume does not publish a ranking of orchestrators for video, and the table above only tells you the limits each vendor states.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume