Script to video AI: turn a finished script into a video
Script to video AI voices your script and cuts a clip to each line. To keep every word, voice it with TTS and cut to its sentence timings.

Script to video AI turns a finished script into a narrated video: the script becomes the voiceover, and each line gets a picture or clip cut to the moment it is spoken. To keep every word exactly, voice the script yourself with text-to-speech and use its sentence timings as the cut points; a video agent can do the whole job, but it may reword the script.
The facts below come from the TTS 1.0 schema in the Sume API reference and the Timeline 1.0, Video Generation, and Create a run docs, read on 2026-09-28. Anything described as current behavior is read from Sume's code.
How do I create an AI video from a script without it being rewritten?
There are two routes, and only one of them keeps the words exactly.
- Voice it yourself. Send the script as the
transcriptof one TTS 1.0 request (POST /v1/tts-1.0/generate). The narration is your text, spoken. Withsegmentation.modeset tosentence(it requirestimestamps.words: true), the result carries gapless sentence segments with their times, which become the cut points in one Timeline 1.0 render. Slideshow with voiceover shows that mapping field by field. - Hand it to an agent. Put the script in the run's
inputand say in the instruction that the narration must use it word for word. On a Format run,inputis carried whole as a file the agent reads and is never truncated, while only about the first 4,000 characters ofinstructionreach the prompt. The agent still decides how to use it, so compare the narration with your script before you publish.
How long will the video from my script be?
As long as the narration. A Timeline render's output always runs for audio.duration_seconds, which you set to the narration's length, so the video ends when the narration does. Text to speech time calculator estimates that length from a word count before you generate anything.
Long scripts hit two limits. One TTS request takes up to 20,000 characters, and audio longer than 1200 s fails with tts_duration_exceeded. A script past either limit needs two or more requests, whose files can be joined as audio.parts[] (up to 20 gapless slices) in one render of up to 1,800 seconds.
How many clips does a script need?
About one per sentence, but a sentence and a clip rarely match in length. Each video model lists its supported_durations in whole seconds at GET /v1/videos/models. A sentence longer than a model's longest clip needs two slots. A clip longer than its sentence is simply cut: the slot plays duration seconds from source_in. A source shorter than its slot is padded or looped, which the render reports as a soft warning, not a failure.
| Step | Limit |
|---|---|
TTS transcript | Up to 20,000 characters per request |
| TTS audio | Over 1200 s fails with tts_duration_exceeded |
| Sentence cut | boundary_lead_ms after the last word: default 70, range 0–500 |
| Generated clip length | The model's supported_durations |
| Render length | audio.duration_seconds: 1–1,800 s |
| Render slots | 1–200 video[] slots, each at least 0.2 s |
| Narration parts | Up to 20 audio.parts[] slices |
What does script-to-video AI not do for you?
Plan around these before you pick the route:
- It doesn't keep each clip's own sound. In current code a Timeline render takes sound only from its audio spine and an optional soundtrack, so a clip's generated audio is dropped and your narration is the only voice.
- It doesn't read local files. Every Timeline URL must already be this workspace's
media.sume.comfile, such as the output of an earlier Sume TTS or video job. - It doesn't caption long videos in one pass. In current code the caption job refuses a source longer than 60 seconds, so caption short cuts or split first.
- It doesn't write the visuals for you on the TTS route. You still choose or generate a picture for each line; Automate faceless short-form videos shows the B-roll step.
Sources
Related posts
More in Use cases
- Security awareness training video: short clips made with AI
A security awareness training video teaches staff one security habit per clip. Make each topic a short captioned AI avatar clip and refresh only what changed.
- Shopify video banner: add a looping video made with AI
A Shopify video banner is a video in a theme section on your home page. Upload the MP4 to Files, pick it in the theme editor, and make the loop with AI.
- Stream starting soon screen: make an animated loop
A stream starting soon screen is a looping video the size of your stream's canvas. How to make an animated one with AI: the loop, the words, the length.
- TikTok Commercial Music Library (CML): what it covers
TikTok's Commercial Music Library is a pre-cleared set of songs businesses may use free, but only on TikTok. What it covers, and music for other uses.
Written by Sume