Video shot list template: columns, example, AI version

A video shot list is one row per shot: number, what we see, framing, move, length, audio, source. A template to copy, plus the columns AI video needs.

5 min readSume
All posts

A video shot list template is a table with one row per shot: the shot number, what the viewer sees, the framing and camera move, how long it lasts, the audio or line spoken over it, and where the footage comes from. Filled in, it tells everyone what to shoot or generate and in what order.

The template is editorial advice. The extra columns for AI-generated shots come from Sume's Video Generation and Timeline 1.0 docs, read on 2026-09-28. It sits between the brief (how to write a video brief) and the edit.

What columns does a shot list need?

Seven columns cover most shoots. Keep each cell short; a shot that needs a paragraph is probably two shots.

  • Shot: a number, so notes and edits can point at it.
  • What we see: subject and action in one line.
  • Framing: wide, medium, close-up, or insert.
  • Move: static, pan, push in, handheld. One move per shot.
  • Length: seconds on screen.
  • Audio: the voiceover line, dialogue, or sound under the shot.
  • Source: filmed, stock, screen recording, or generated.

Is there a shot list template I can copy?

Paste this into a sheet or a doc. The example is a 20-second vertical product video with a voiceover.

SHOT | WHAT WE SEE                        | FRAMING  | MOVE      | SEC | AUDIO / LINE                     | SOURCE
1    | Hand lifts a bottle off a desk     | Close-up | Static    | 4   | VO: "Your morning, in one bottle." | Generated
2    | Bottle on a sunny kitchen counter  | Medium   | Push in   | 5   | VO: "Cold brew, no sugar."        | Generated
3    | Person drinks, smiles              | Medium   | Handheld  | 5   | VO: "Ready in the fridge."        | Filmed
4    | Logo and price on a plain backdrop | Wide     | Static    | 6   | VO: "Order today."                | Generated

What changes when the shots are AI generated?

Each generated row becomes one generation request, so the row has to hold what the request needs. On Sume's POST /v1/videos, that is a prompt (your What we see, Framing, and Move cells written as one sentence), a duration, an aspect_ratio, and optionally a start image sent in frame_images as the first_frame. Reusing one start image across rows is covered in consistent character across AI video shots.

The limits are per model, and they are not uniform, so check the catalog before you fix the Length column. GET /v1/videos/models lists what each model accepts:

From Video Generation, read 2026-09-28.
Shot list columnRequest fieldCheck against
What we see, Framing, MovepromptRequired: a text description of the video.
Lengthdurationsupported_durations, in whole seconds.
Shape (one per list)aspect_ratiosupported_aspect_ratios, such as 9:16 or 16:9.
Start imageframe_imagessupported_frame_images: first_frame, last_frame.
Sound in the clipgenerate_audiogenerate_audio on the model: whether it can make a track.

How does the shot list become the edit?

The row order and the Length column are the edit decision. In Sume's Timeline 1.0 render, each generated row becomes one slot in video[], with the finished clip as its source_url and the Length cell as its duration. The slot rules are covered in storyboard to video with AI, and what is an edit decision list maps the same fields onto EDL columns. What the shot list adds:

  • A running Start column: row 1 starts at 0, and each later row's start is the sum of the lengths above it. That is the slot's start, which must increase from row to row.
  • Put the Audio column into one voiceover track and send it as the render's audio spine. In current code a Timeline render takes sound only from that spine and an optional soundtrack, so each clip's own audio is dropped.
  • Timeline reads only this workspace's media.sume.com files, such as the outputs of your earlier Sume generations. Filmed rows have to be joined in your own editor.

Can an AI agent write the shot list for me?

Yes, if you give it the brief and review the list before anything is generated. Ask for the plan as structured data (one object per shot, with the fields above) and read it the way you would read a person's draft: count the shots, add up the lengths, and check every length against the model's list. Structured output is covered in JSON schema templates for video runs. For the prompt inside each row, see how to write an AI video prompt.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume