Add an AI voiceover to a silent video: TTS, then a timeline render

Add a voiceover to a silent clip with Sume: one TTS job for the narration, then one Timeline 1.0 render that lays the audio over your video.

4 min readSume
All posts

To add an AI voiceover to a silent video, make the narration first and render second. Sume's TTS 1.0 turns your script into an audio file. Timeline 1.0 then takes that file as its audio spine and places your clip on top, and it returns one MP4. The silent clip is never edited in place, so you can re-render with a new take whenever the script changes.

Step 1: make the narration

Both jobs follow the same lifecycle. You submit, the response carries a job id, and you read GET /v1/jobs/:id/status and then /result. The default mode is async. mode: "sync" waits up to 30 seconds and falls back to a 202 if the job is not done, so a slow job is polled and not resubmitted.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: vo-silent-001" \
  -d "{\"transcript\": \"Three steps, no setup.\", \"voice\": {\"id\": \"$VOICE_ID\"}, \"output_format\": {\"container\": \"wav\", \"sample_rate\": 44100, \"encoding\": \"pcm_s16le\"}}"

Step 2: render the timeline

The result holds the audio file. transcript is limited to 20000 characters, and TTS 1.0 audio longer than 1200 seconds fails with tts_duration_exceeded. Set language for any non-English script. Take the voice from a workspace avatar (avatar_handle) or pass voice.id as above.

Timeline 1.0 needs every URL to be your workspace's media.sume.com artifact, so import the silent clip first with POST /v1/media-imports. Then send audio.url and audio.duration_seconds (1 to 1800) plus one or more video[] slots. If the clip is shorter than the audio, the render pads or loops it and reports a soft warning, which is not a failure. You can run POST /v1/timeline-1.0/plan first. It is unbilled and returns the cost estimate before you commit.

What you pay and get

What each step costs and returns, from the Sume docs and the API schema (read 2026-10-06):

Voiceover pipeline, read 2026-10-06
StepEndpointBilling basisLimit that matters
NarrationPOST /v1/tts-1.0/generatePer character20000 characters, 1200 s of audio
Import clipPOST /v1/media-importsSee the API referenceTimeline needs media.sume.com URLs
PreflightPOST /v1/timeline-1.0/planUnbilledReturns estimated cost
RenderPOST /v1/timeline-1.0/render$0.10 per output minute, rounded upAudio 1 to 1800 s, 1 to 200 video slots

Mistakes that cost a second render

Most reruns come from length, not from voice. Read the audio duration from the TTS result and use that number as audio.duration_seconds, so the picture and the voice end together.

Write the script to the clip, not the clip to the script. Count the characters first, since TTS is billed per character and the text limit is 20000. Keep one idea per sentence so a later retake replaces one line and not the whole read. If you retake a line, send a new idempotency key, because the same key with the same body is read as a retry.

The render default is 1080 by 1920. Both writes need an Idempotency-Key, so a retried request returns the same job and not a second charge. Confirm the live rate in GET /v1/catalog.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume