Script to video AI: turn a finished script into a video

Script to video AI voices your script and cuts a clip to each line. To keep every word, voice it with TTS and cut to its sentence timings.

5 min readSume
All posts

Script to video AI turns a finished script into a narrated video: the script becomes the voiceover, and each line gets a picture or clip cut to the moment it is spoken. To keep every word exactly, voice the script yourself with text-to-speech and use its sentence timings as the cut points; a video agent can do the whole job, but it may reword the script.

The facts below come from the TTS 1.0 schema in the Sume API reference and the Timeline 1.0, Video Generation, and Create a run docs, read on 2026-09-28. Anything described as current behavior is read from Sume's code.

How do I create an AI video from a script without it being rewritten?

There are two routes, and only one of them keeps the words exactly.

  • Voice it yourself. Send the script as the transcript of one TTS 1.0 request (POST /v1/tts-1.0/generate). The narration is your text, spoken. With segmentation.mode set to sentence (it requires timestamps.words: true), the result carries gapless sentence segments with their times, which become the cut points in one Timeline 1.0 render. Slideshow with voiceover shows that mapping field by field.
  • Hand it to an agent. Put the script in the run's input and say in the instruction that the narration must use it word for word. On a Format run, input is carried whole as a file the agent reads and is never truncated, while only about the first 4,000 characters of instruction reach the prompt. The agent still decides how to use it, so compare the narration with your script before you publish.

How long will the video from my script be?

As long as the narration. A Timeline render's output always runs for audio.duration_seconds, which you set to the narration's length, so the video ends when the narration does. Text to speech time calculator estimates that length from a word count before you generate anything.

Long scripts hit two limits. One TTS request takes up to 20,000 characters, and audio longer than 1200 s fails with tts_duration_exceeded. A script past either limit needs two or more requests, whose files can be joined as audio.parts[] (up to 20 gapless slices) in one render of up to 1,800 seconds.

How many clips does a script need?

About one per sentence, but a sentence and a clip rarely match in length. Each video model lists its supported_durations in whole seconds at GET /v1/videos/models. A sentence longer than a model's longest clip needs two slots. A clip longer than its sentence is simply cut: the slot plays duration seconds from source_in. A source shorter than its slot is padded or looped, which the render reports as a soft warning, not a failure.

From the Sume API reference, Timeline 1.0, and Video Generation, read 2026-09-28.
StepLimit
TTS transcriptUp to 20,000 characters per request
TTS audioOver 1200 s fails with tts_duration_exceeded
Sentence cutboundary_lead_ms after the last word: default 70, range 0–500
Generated clip lengthThe model's supported_durations
Render lengthaudio.duration_seconds: 1–1,800 s
Render slots1–200 video[] slots, each at least 0.2 s
Narration partsUp to 20 audio.parts[] slices

What does script-to-video AI not do for you?

Plan around these before you pick the route:

  • It doesn't keep each clip's own sound. In current code a Timeline render takes sound only from its audio spine and an optional soundtrack, so a clip's generated audio is dropped and your narration is the only voice.
  • It doesn't read local files. Every Timeline URL must already be this workspace's media.sume.com file, such as the output of an earlier Sume TTS or video job.
  • It doesn't caption long videos in one pass. In current code the caption job refuses a source longer than 60 seconds, so caption short cuts or split first.
  • It doesn't write the visuals for you on the TTS route. You still choose or generate a picture for each line; Automate faceless short-form videos shows the B-roll step.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume