How to make an AI documentary video, chapter by chapter
Make an AI documentary video from a researched script: voice it in chapters, make a shot per beat, and assemble up to 30 minutes in one render.

To make an AI documentary video, write a researched script in chapters, voice each chapter with text to speech, make one short visual for each beat of the narration, and assemble the visuals over the narration in one long render. AI does the production; the facts come from your research, and the generated pictures are illustrations, not footage of real events.
A documentary is long narration over many short shots, so the limits that matter are per-job lengths. The Sume numbers below come from the Timeline 1.0, Timeline audio, and Video generation docs and the TTS schema in the Sume API reference, read on 2026-09-28.
How long can an AI documentary be?
Up to 30 minutes per render: Timeline 1.0 takes an output length of 1–1,800 seconds. Every other step has a smaller unit, so a long film is built from many jobs:
| Step | Limit per job | What it means for a documentary |
|---|---|---|
| Narration (TTS 1.0) | 20,000 characters of transcript; audio over 1,200 s fails | Voice one chapter per job |
| Joining narration (Timeline audio) | 1–20 parts; output ≤ 1,800 s | One gapless narration file from up to 20 chapters |
Shots (POST /v1/videos) | Model-listed durations, such as 4–15 s on seedance-2 or 4–30 s on seedance-2.5 | Several shots per chapter |
| Assembly (Timeline 1.0) | 1–1,800 s; 1–200 video[] slots | 200 slots over 30 minutes averages 9 s per shot |
Captions (POST /v1/video-captions) | 60 s per job (current code) | Caption in parts |
How do I voice a long narration?
Split the script at chapter breaks and send each chapter to POST /v1/tts-1.0/generate as transcript, in the same voice each time: one avatar_id or avatar_handle, or one voice.id. Add timestamps: { "words": true } to get word start and end times, which tell you where each sentence lands for cutting shots.
Then join the chapters. POST /v1/timeline-1.0/audio with operation: "concat" joins up to 20 Sume-hosted parts into one gapless file, with no re-synthesis and no silence at the seams. Its result lists segments[] with each part's start, the offsets to place each chapter's shots against. For a join used only inside one render, put the chapters on the render's audio.parts[] instead.
Where do the visuals come from?
From one generated clip per beat. Write a shot list from the script, one line per visual idea, and generate each with POST /v1/videos: a prompt describing the scene, aspect_ratio: "16:9" for a landscape film, and a duration the model lists. To open a shot on a specific picture, send it as a first_frame in frame_images.
A still placed straight in a Timeline slot is held static, and any motion on it is ignored, so a shot that should move has to be generated as a clip. Label generated scenes as illustrations wherever a viewer could mistake them for archive footage.
How do I assemble and caption the documentary?
One Timeline 1.0 render lays every shot on the narration. Set output.width and output.height (even integers 256–2160) for 16:9, since the default is 1080×1920 vertical. The output's sound is the narration spine plus an optional soundtrack bed, so any sound you want must be on one of those two. Every URL must be a Sume-hosted file, such as an earlier Sume job's output. The sample is a 17-second test cut with two shots; the full film repeats the pattern for up to 200 slots. How to assemble a long-form video covers every field.
In current code a caption job refuses a video over 60 seconds, so a long documentary is cut into short parts, captioned, and rejoined: Add captions to a long video shows the steps.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: documentary-render-001" \
-d '{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/narration.wav", "duration_seconds": 17 },
"output": { "width": 1920, "height": 1080 },
"soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/score.mp3", "loop": true, "duck_db": 10 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/shot-001.mp4", "start": 0, "duration": 9 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/shot-002.mp4", "start": 9, "duration": 8,
"transition": { "type": "dissolve", "duration": 0.5 } }
]
}'What does an AI documentary cost?
Each step bills on its own. Narration is $0.0475 per 1,000 characters, and spaces and punctuation count. The render is $0.10 per output minute, reserved as ceil(audio.duration_seconds / 60) minutes. Generated clips are reserved at the provider's list price × 1.25, and a pinned model's rate is on GET /v1/videos/models. The rates are on API pricing, plus a 5.5% agent fee by default. Check a render for free first: POST /v1/timeline-1.0/plan returns its length, slot count, and estimated cost without creating a job.
Sources
Related posts
More in Use cases
- How to make an AI voiceover: from script to audio file
Make an AI voiceover in five steps: write the script, pick a voice, generate the speech, check its length, and download the file. The Sume API way.
- How to make an unboxing video with AI, step by step
An unboxing video shows hands opening the package and revealing the product. How to make one with AI from a box photo and a product photo.
- How to create meeting minutes from an audio recording
Create meeting minutes from a recording in two steps: transcribe the audio, then have a language model draft the summary, decisions, and action items.
- Create motivational videos with AI: voice, shots, and music
Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.
Written by Sume