Add AI voiceover in another language to a silent screen recording

Write the narration, generate it with a language code, lay it over the recording in a Timeline render, then caption it. What each step bills and what it limits.

5 min readSume
All posts

To add a voiceover in another language to a silent screen recording, import the recording, write the narration in the target language, generate it with TTS 1.0 and a language code, and lay the audio over the video with a Timeline 1.0 render. If you want captions after that, run the finished video through the caption job, because by then it has speech in it.

This follows the Sume docs for Timeline 1.0 and video captions, read 2026-10-03.

Which step needs what?

The render's audio length decides the output length. If your narration is shorter than the recording, the recording is cut to match; if it is longer, the video slot needs to cover it, since coverage may trail the audio by at most half a second.

  • Import the recording first. Timeline and caption calls take a media.sume.com URL from your workspace, and an off-host URL is rejected.
  • Write the narration in the target language. Sume has no translation call, so the script is yours.
  • Generate speech with POST /v1/tts-1.0/generate, setting language and a voice for that language.
  • Render with Timeline 1.0: the speech as audio.url, the recording as a video[] slot.
  • Caption the result if you want burned text.

How do you fit a landscape recording?

A desktop recording is landscape and the render default is 1080x1920, a vertical frame. Set output.width and output.height to even numbers between 256 and 2160, for example 1920 by 1080, and choose fit: cover crops to fill, contain keeps the whole frame, blur fills the bars with a blurred copy. For a screen recording contain keeps every pixel of the interface, which is usually what you want.

You can check the document first with the unbilled plan endpoint, which returns the duration and the billable minutes without creating a job.

curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio": { "url": "'"$VOICE_URL"'", "duration_seconds": 42 },
    "video": [{
      "source_url": "'"$SCREEN_URL"'",
      "start": 0, "duration": 42, "fit": "contain"
    }],
    "output": { "width": 1920, "height": 1080 }
  }'

What does each step cost?

Voiceover on a silent recording, Sume rates, read 2026-10-03
StepRateNotes
TTS 1.0$0.0475 per 1,000 characters1 cent minimum per job
Timeline render$0.10 per output minute, rounded upA 42-second video bills as 1 minute
Burned captions$0.20 per video up to 60 secondsNeeds speech or authored cues
Plan checkFreeDoes not create a job

Why do captions fail on the silent original?

The caption job transcribes the clip's audio, so a silent recording fails with caption_no_speech, and the error suggests authored overlay captions instead. Two routes work. Either caption the finished video that now carries your voiceover, where language is a speech-to-text hint for what was spoken, or pass your own cues with text and start and end times to burn exact wording without transcription.

The second route lets you caption in a language other than the voice, or put short labels over the interface at chosen moments. Each cue is text up to 400 characters, with times within 60 seconds, and a job takes up to 200 cues.

How do you match the words to the screen?

Write the narration as short sentences, one per action on screen. Ask TTS 1.0 for timestamps.words and sentence segmentation (which needs a wav or raw output) so each sentence comes back as its own slice with its own duration. Then join the slices with Timeline audio concat, which returns each slice's offset in the combined file.

Those offsets are the numbers to re-base your video slots against: slot one starts at 0, slot two at the offset of sentence two, and so on. Trim the recording into scenes with video trim first if one long take needs to change at each sentence.

Do a dry pass before the paid one: render the plan check, read billable_minutes, and confirm the video slot durations add up to the audio length. A mismatch there is a free error, while the same mistake in a paid render costs a whole rounded-up minute and a wait.

If the narration for a step is longer than that step on screen, you can hold a still frame or lengthen the slot, but Sume documents no way to slow the recording itself. Rewriting the sentence shorter is cheaper than fighting the timing.

What Sume will not do for you

Sume does not watch the recording and write the narration, does not translate it, and does not time the speech to the cursor movements. Matching words to what is on screen is done by you: split the narration into sentences, generate them with sentence segmentation, and use the segment offsets to set each video[].start. For the plain English version of this job, see add voiceover to video.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume