Microdrama script format: cliffhanger and AI shot list

Write a microdrama as a beat sheet, then turn each beat into one AI shot no longer than the model allows. A format with the cliffhanger as the last frame.

4 min readSume
All posts

Write a microdrama script as a beat sheet with one line per shot, ending on the cliffhanger beat, then give every line a length no longer than the video model accepts. With AI video the script and the shot list are the same document, because each line becomes one generation request.

The limits that shape it come from the Video generation docs: duration ceilings differ by model, and a vertical frame is aspect_ratio: "9:16". Story advice below is craft, not a Sume rule.

What does a beat sheet for an AI microdrama look like?

Give each beat a number, a location, one action, one line of dialogue at most, and a length. Keep the prompt for each beat under one sentence of action. Long, multi-action prompts are where the shot stops matching your intent.

Example episode beat sheet, with model limits from the Video generation docs, read 2026-09-29.
BeatJob in the storyShot length rule
1 HookThe problem lands in the first imageStay within the model's supported_durations
2 TurnA message, a knock, a revealgemini-omni-flash-1.1 accepts 3–10 seconds
3 EscalateThe lead has to chooseseedance-2.5 accepts 4–30 seconds
4 CliffhangerEnd on the question, not the answerSet a last_frame so the shot ends where you planned

How do I write the cliffhanger so the model can end on it?

Decide the final image before you write the shot. A frame you already hold, such as the lead staring at a phone, can be sent as frame_images with frame_type: last_frame; the docs list first_frame and last_frame as the accepted values, and a model only accepts them if its supported_frame_images lists them. The next episode's first shot can open on that same image as its first_frame, which is a simple way to make two episodes join.

Check the model's entry in GET /v1/videos/models first. If it does not list last_frame, end the shot with a plain description of the closing image in the prompt and choose the take you like.

How long can each shot be?

Limits are per model. The docs say seedance-2.5 accepts 4–30 seconds at 480p/720p/1080p, wan-3.0 accepts 2–30 seconds, and every other catalog model is capped at 15 seconds, while gemini-omni-flash-1.1 is 3–10 seconds in 16:9 or 9:16. The Video Router docs tell you to read capabilities from GET /v1/video-router/models rather than assume one envelope. Plan shots at the shortest limit among the models you might try, so you can swap models without rewriting the sheet.

What does one line of the shot list become?

One request per beat. Use a stable Idempotency-Key per shot (ep03-beat02) so a retry after a dropped connection returns the original job instead of a second job.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ep03-beat02" \
  -d '{
    "model": "seedance-2.5",
    "prompt": "A woman reads a text at a night bus stop, then looks up sharply",
    "resolution": "720p",
    "duration": 8,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

How do I join the beats into one episode?

Sume's Timeline 1.0 takes one audio spine plus ordered video slots and assembles a single MP4, so the beats render as separate shots and join in one step. For keeping the same face across those shots, see keeping characters consistent across episodes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume