Which LLM for each step of a video agent: script to assembly
A hypothesis for matching LLMs to video agent steps (script, shot list, frame check, assembly), using documented tool prices, to test with a Format A/B run.

A sensible starting split is a stronger model for the script and shot list, an image-capable model for frame checks, and the cheapest reliable model for assembly. That is a hypothesis, not a finding. Nothing here has been benchmarked, and the way to settle it is to run your own input through two models and read the receipts.
What the docs do support is the cost of the tools around the model, which shows where model choice matters and where it does not.
What do the tools cost?
Assembly is cheap in tool terms, so a wrong assembly call costs more in retries and LLM turns than in tool fees. That argues for a model that follows a schema reliably, not one that reasons the longest.
| Step | Tool | Documented price |
|---|---|---|
| Probe a clip | video_inspect | Probe and stills unbilled; audio transcript $0.01 per minute |
| Pull stills | video_frames | Unbilled per the docs |
| Cut a clip | video-trim | $0.02 per job |
| Assemble a cut | Timeline 1.0 | $0.10 per output minute, rounded up |
| Plan a Timeline | Timeline plan | Unbilled compile preflight |
Which model for the script and shot list?
This is the step where one good plan saves renders, so it is the step to spend on. Try your strongest affordable model at its default effort first, and raise effort only for this turn. Opus 5.5 against Fable 5.1 is the comparison to run if you are already on Claude. Do not assume the bigger model writes better ad copy; test with two.
Which model for the frame check?
The agent has to see stills, so the model must take image input. Check the model's vendor page or catalog row for image input before you assign it this step, since not every text model reads images. Video frames pulls stills without a per-call charge, so the cost is the LLM turn that reads them. Keep the number of frames small and ask for a structured pass or fail, so the turn stays short.
Which model for assembly?
Assembly is mostly writing a valid Timeline and checking the unbilled plan before paying for the render. A cheaper model at low effort is a fair first try, since the plan step catches a bad Timeline before it costs anything. If it fails the plan twice, move up a tier rather than looping.
How do you test the split?
The Formats model field picks one orchestrator per run, so you cannot assign a different LLM to each step inside one run from the API. Run the whole Format on each candidate, as in the A/B recipe, and read usage.debited_usd_micros and the output. If one model wins on plan quality and another on cost, split the work yourself: plan with one Format run, then continue or start a second run on the cheaper model.
What would change the split?
Three things would move it. First, a failure pattern: if the cheap model keeps producing Timelines that fail the plan check, the saving is false and assembly belongs on the stronger model. Second, price changes at the vendor, such as a tier boundary at a long context, which makes long threads costlier; the Grok 4.7 200k price jump is one case. Third, the effort setting, which can make a mid-priced model do a top-priced model's job on one step only.
A fourth thing is volume. At a handful of videos a month the whole debate is worth a few dollars, and the time is better spent on the brief. At thousands of runs a month a ten-cent difference per run is real money, and a proper A/B pays for itself in a day.
Keep the experiment small. Pick one Format, one input and two candidates, run each three times, and write down cost, completion and a quality score you define in advance. The result is an answer for your workload, which is the only kind of answer this question has.
Sources
Related posts
More in Models
- LTX-2.3 LoRAs on LTX-2.5: most run unchanged; Sume has no LoRA field
Lightricks says the large majority of LTX-2.3 LoRAs and IC-LoRAs run on LTX-2.5 unchanged, with a few exceptions to test. Sume has no LoRA field.
- LTX-2.5 text encoder: use the bundled Gemma 4 12B, not stock
LTX-2.5 needs its own Gemma 4 12B text encoder; Google's stock Gemma 4 release is not a substitute. What to download and what the 66 GB package holds.
- LTX-2.5 distilled or dev transformer: which file to download
Download the distilled bf16 transformer to generate (fixed 8-step schedule, CFG 1) and the dev one to train. int8 files are ComfyUI-only. Sume lists no LTX id.
- LTX-2.5 duration predictor: optional, and Sume sets seconds
LTX-2.5 has an optional duration predictor that sets the frame count from the prompt. Sume has no such mode: you send duration in whole seconds.
Written by Sume