Video shot list template: columns, example, AI version
A video shot list is one row per shot: number, what we see, framing, move, length, audio, source. A template to copy, plus the columns AI video needs.

A video shot list template is a table with one row per shot: the shot number, what the viewer sees, the framing and camera move, how long it lasts, the audio or line spoken over it, and where the footage comes from. Filled in, it tells everyone what to shoot or generate and in what order.
The template is editorial advice. The extra columns for AI-generated shots come from Sume's Video Generation and Timeline 1.0 docs, read on 2026-09-28. It sits between the brief (how to write a video brief) and the edit.
What columns does a shot list need?
Seven columns cover most shoots. Keep each cell short; a shot that needs a paragraph is probably two shots.
- Shot: a number, so notes and edits can point at it.
- What we see: subject and action in one line.
- Framing: wide, medium, close-up, or insert.
- Move: static, pan, push in, handheld. One move per shot.
- Length: seconds on screen.
- Audio: the voiceover line, dialogue, or sound under the shot.
- Source: filmed, stock, screen recording, or generated.
Is there a shot list template I can copy?
Paste this into a sheet or a doc. The example is a 20-second vertical product video with a voiceover.
SHOT | WHAT WE SEE | FRAMING | MOVE | SEC | AUDIO / LINE | SOURCE
1 | Hand lifts a bottle off a desk | Close-up | Static | 4 | VO: "Your morning, in one bottle." | Generated
2 | Bottle on a sunny kitchen counter | Medium | Push in | 5 | VO: "Cold brew, no sugar." | Generated
3 | Person drinks, smiles | Medium | Handheld | 5 | VO: "Ready in the fridge." | Filmed
4 | Logo and price on a plain backdrop | Wide | Static | 6 | VO: "Order today." | GeneratedWhat changes when the shots are AI generated?
Each generated row becomes one generation request, so the row has to hold what the request needs. On Sume's POST /v1/videos, that is a prompt (your What we see, Framing, and Move cells written as one sentence), a duration, an aspect_ratio, and optionally a start image sent in frame_images as the first_frame. Reusing one start image across rows is covered in consistent character across AI video shots.
The limits are per model, and they are not uniform, so check the catalog before you fix the Length column. GET /v1/videos/models lists what each model accepts:
| Shot list column | Request field | Check against |
|---|---|---|
| What we see, Framing, Move | prompt | Required: a text description of the video. |
| Length | duration | supported_durations, in whole seconds. |
| Shape (one per list) | aspect_ratio | supported_aspect_ratios, such as 9:16 or 16:9. |
| Start image | frame_images | supported_frame_images: first_frame, last_frame. |
| Sound in the clip | generate_audio | generate_audio on the model: whether it can make a track. |
How does the shot list become the edit?
The row order and the Length column are the edit decision. In Sume's Timeline 1.0 render, each generated row becomes one slot in video[], with the finished clip as its source_url and the Length cell as its duration. The slot rules are covered in storyboard to video with AI, and what is an edit decision list maps the same fields onto EDL columns. What the shot list adds:
- A running Start column: row 1 starts at 0, and each later row's start is the sum of the lengths above it. That is the slot's
start, which must increase from row to row. - Put the Audio column into one voiceover track and send it as the render's
audiospine. In current code a Timeline render takes sound only from that spine and an optional soundtrack, so each clip's own audio is dropped. - Timeline reads only this workspace's
media.sume.comfiles, such as the outputs of your earlier Sume generations. Filmed rows have to be joined in your own editor.
Can an AI agent write the shot list for me?
Yes, if you give it the brief and review the list before anything is generated. Ask for the plan as structured data (one object per shot, with the fields above) and read it the way you would read a person's draft: count the shots, add up the lengths, and check every length against the model's list. Structured output is covered in JSON schema templates for video runs. For the prompt inside each row, see how to write an AI video prompt.
Sources
Related posts
More in Agents
- Video agent API: turn a brief into a finished, edited video
Yes: a video agent API turns a brief into a finished, edited video. What Sume's Agent Completions and Formats take, return, cost, and won't do.
- What is an MCP server? A plain definition with examples
An MCP server is a program that gives AI apps tools, data, and prompt templates through the Model Context Protocol. How it works, with examples.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
Written by Sume