Add AI voiceover in another language to a silent screen recording
Write the narration, generate it with a language code, lay it over the recording in a Timeline render, then caption it. What each step bills and what it limits.

To add a voiceover in another language to a silent screen recording, import the recording, write the narration in the target language, generate it with TTS 1.0 and a language code, and lay the audio over the video with a Timeline 1.0 render. If you want captions after that, run the finished video through the caption job, because by then it has speech in it.
This follows the Sume docs for Timeline 1.0 and video captions, read 2026-10-03.
Which step needs what?
The render's audio length decides the output length. If your narration is shorter than the recording, the recording is cut to match; if it is longer, the video slot needs to cover it, since coverage may trail the audio by at most half a second.
- Import the recording first. Timeline and caption calls take a
media.sume.comURL from your workspace, and an off-host URL is rejected. - Write the narration in the target language. Sume has no translation call, so the script is yours.
- Generate speech with
POST /v1/tts-1.0/generate, settinglanguageand a voice for that language. - Render with Timeline 1.0: the speech as
audio.url, the recording as avideo[]slot. - Caption the result if you want burned text.
How do you fit a landscape recording?
A desktop recording is landscape and the render default is 1080x1920, a vertical frame. Set output.width and output.height to even numbers between 256 and 2160, for example 1920 by 1080, and choose fit: cover crops to fill, contain keeps the whole frame, blur fills the bars with a blurred copy. For a screen recording contain keeps every pixel of the interface, which is usually what you want.
You can check the document first with the unbilled plan endpoint, which returns the duration and the billable minutes without creating a job.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": { "url": "'"$VOICE_URL"'", "duration_seconds": 42 },
"video": [{
"source_url": "'"$SCREEN_URL"'",
"start": 0, "duration": 42, "fit": "contain"
}],
"output": { "width": 1920, "height": 1080 }
}'What does each step cost?
| Step | Rate | Notes |
|---|---|---|
| TTS 1.0 | $0.0475 per 1,000 characters | 1 cent minimum per job |
| Timeline render | $0.10 per output minute, rounded up | A 42-second video bills as 1 minute |
| Burned captions | $0.20 per video up to 60 seconds | Needs speech or authored cues |
| Plan check | Free | Does not create a job |
Why do captions fail on the silent original?
The caption job transcribes the clip's audio, so a silent recording fails with caption_no_speech, and the error suggests authored overlay captions instead. Two routes work. Either caption the finished video that now carries your voiceover, where language is a speech-to-text hint for what was spoken, or pass your own cues with text and start and end times to burn exact wording without transcription.
The second route lets you caption in a language other than the voice, or put short labels over the interface at chosen moments. Each cue is text up to 400 characters, with times within 60 seconds, and a job takes up to 200 cues.
How do you match the words to the screen?
Write the narration as short sentences, one per action on screen. Ask TTS 1.0 for timestamps.words and sentence segmentation (which needs a wav or raw output) so each sentence comes back as its own slice with its own duration. Then join the slices with Timeline audio concat, which returns each slice's offset in the combined file.
Those offsets are the numbers to re-base your video slots against: slot one starts at 0, slot two at the offset of sentence two, and so on. Trim the recording into scenes with video trim first if one long take needs to change at each sentence.
Do a dry pass before the paid one: render the plan check, read billable_minutes, and confirm the video slot durations add up to the audio length. A mismatch there is a free error, while the same mistake in a paid render costs a whole rounded-up minute and a wait.
If the narration for a step is longer than that step on screen, you can hold a still frame or lengthen the slot, but Sume documents no way to slow the recording itself. Rewriting the sentence shorter is cheaper than fighting the timing.
What Sume will not do for you
Sume does not watch the recording and write the narration, does not translate it, and does not time the speech to the cursor movements. Matching words to what is on screen is done by you: split the narration into sentences, generate them with sentence segmentation, and use the segment offsets to set each video[].start. For the plain English version of this job, see add voiceover to video.
Sources
Related posts
More in Use cases
- Affiliate link on Instagram: does it need the paid partnership label?
Meta says content with an affiliate link should carry the Paid partnership label, even with no business partner tagged. What that means for AI clips.
- AI avatar mock interview: live interviewer or question clips
Tavus lists an Interviewer PAL as a live use case. Sume can render each interview question as a short avatar clip with a thinking pause. Where each one fits.
- AI avatar of a real person: release checklist before the photo
Before you turn an employee's or creator's photo into a Sume avatar, get a written release. A checklist for scope, term and revocation, matched to the API.
- AI avatar sales agent: live SDR or personalised clips?
Tavus builds live SDR avatars on a per-minute plan. Sume renders one 4-60 second avatar clip per lead from a script. How to pick, plus a Python loop.
Written by Sume