How to make an AI documentary video, chapter by chapter

Make an AI documentary video from a researched script: voice it in chapters, make a shot per beat, and assemble up to 30 minutes in one render.

5 min readSume
All posts

To make an AI documentary video, write a researched script in chapters, voice each chapter with text to speech, make one short visual for each beat of the narration, and assemble the visuals over the narration in one long render. AI does the production; the facts come from your research, and the generated pictures are illustrations, not footage of real events.

A documentary is long narration over many short shots, so the limits that matter are per-job lengths. The Sume numbers below come from the Timeline 1.0, Timeline audio, and Video generation docs and the TTS schema in the Sume API reference, read on 2026-09-28.

How long can an AI documentary be?

Up to 30 minutes per render: Timeline 1.0 takes an output length of 1–1,800 seconds. Every other step has a smaller unit, so a long film is built from many jobs:

From Timeline 1.0, Timeline audio, Video generation, and the Sume API reference, read 2026-09-28.
StepLimit per jobWhat it means for a documentary
Narration (TTS 1.0)20,000 characters of transcript; audio over 1,200 s failsVoice one chapter per job
Joining narration (Timeline audio)1–20 parts; output ≤ 1,800 sOne gapless narration file from up to 20 chapters
Shots (POST /v1/videos)Model-listed durations, such as 4–15 s on seedance-2 or 4–30 s on seedance-2.5Several shots per chapter
Assembly (Timeline 1.0)1–1,800 s; 1–200 video[] slots200 slots over 30 minutes averages 9 s per shot
Captions (POST /v1/video-captions)60 s per job (current code)Caption in parts

How do I voice a long narration?

Split the script at chapter breaks and send each chapter to POST /v1/tts-1.0/generate as transcript, in the same voice each time: one avatar_id or avatar_handle, or one voice.id. Add timestamps: { "words": true } to get word start and end times, which tell you where each sentence lands for cutting shots.

Then join the chapters. POST /v1/timeline-1.0/audio with operation: "concat" joins up to 20 Sume-hosted parts into one gapless file, with no re-synthesis and no silence at the seams. Its result lists segments[] with each part's start, the offsets to place each chapter's shots against. For a join used only inside one render, put the chapters on the render's audio.parts[] instead.

Where do the visuals come from?

From one generated clip per beat. Write a shot list from the script, one line per visual idea, and generate each with POST /v1/videos: a prompt describing the scene, aspect_ratio: "16:9" for a landscape film, and a duration the model lists. To open a shot on a specific picture, send it as a first_frame in frame_images.

A still placed straight in a Timeline slot is held static, and any motion on it is ignored, so a shot that should move has to be generated as a clip. Label generated scenes as illustrations wherever a viewer could mistake them for archive footage.

How do I assemble and caption the documentary?

One Timeline 1.0 render lays every shot on the narration. Set output.width and output.height (even integers 256–2160) for 16:9, since the default is 1080×1920 vertical. The output's sound is the narration spine plus an optional soundtrack bed, so any sound you want must be on one of those two. Every URL must be a Sume-hosted file, such as an earlier Sume job's output. The sample is a 17-second test cut with two shots; the full film repeats the pattern for up to 200 slots. How to assemble a long-form video covers every field.

In current code a caption job refuses a video over 60 seconds, so a long documentary is cut into short parts, captioned, and rejoined: Add captions to a long video shows the steps.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: documentary-render-001" \
  -d '{
    "audio": { "url": "https://media.sume.com/artifacts/artf_demo/narration.wav", "duration_seconds": 17 },
    "output": { "width": 1920, "height": 1080 },
    "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/score.mp3", "loop": true, "duck_db": 10 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/shot-001.mp4", "start": 0, "duration": 9 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/shot-002.mp4", "start": 9, "duration": 8,
        "transition": { "type": "dissolve", "duration": 0.5 } }
    ]
  }'

What does an AI documentary cost?

Each step bills on its own. Narration is $0.0475 per 1,000 characters, and spaces and punctuation count. The render is $0.10 per output minute, reserved as ceil(audio.duration_seconds / 60) minutes. Generated clips are reserved at the provider's list price × 1.25, and a pinned model's rate is on GET /v1/videos/models. The rates are on API pricing, plus a 5.5% agent fee by default. Check a render for free first: POST /v1/timeline-1.0/plan returns its length, slot count, and estimated cost without creating a job.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume