Replace the voice in a finished video: detach, transcribe, TTS, render
Four Sume calls swap the narration in a finished video: detach the audio, transcribe it, voice your edited script, render. About 17 cents for one minute.

To replace the narration in a finished video on Sume, chain four jobs: detach the audio from the video, transcribe it to get the script, voice your edited script with TTS, then render the video with the new audio as the spine. For a one-minute video with a 1,000-character script the rates add up to $0.1675: $0.01 detach, $0.01 transcription, $0.0475 TTS and $0.10 for one rendered minute (API reference).
Step by step
- Detach:
POST /v1/audio-detachtakes avideo_urlon your workspace'smedia.sume.com. Import the file first if it lives elsewhere. Ask forchannels: "mono"andsample_rate: 16000, which the detach docs call the speech-to-text shape. - Transcribe:
POST /v1/stt-1.0/transcribewith the detachedaudio_url, up to 600 seconds per job. Edit the returnedtext. - Voice it:
POST /v1/tts-1.0/generatewith your edited text as a wav, so the render step receives a clean file. - Render:
POST /v1/timeline-1.0/renderwithaudio.urlset to the TTS file,audio.duration_secondsset to its length, and the original clip as the singlevideo[]slot (timeline docs).
The render step
The sound track of a Sume render comes from the audio you declare, plus an optional soundtrack, not from the clip. Still, listen to the output to confirm. If the new narration is longer than the clip, the render can pad or loop the video and says so in a warning, so keep the edited script close to the original length.
{
"audio": {"url": "https://media.sume.com/artifacts/artf_new/voice.wav",
"duration_seconds": 58},
"video": [{"source_url": "https://media.sume.com/artifacts/artf_old/clip.mp4",
"start": 0, "duration": 58, "fit": "cover"}]
}Check the length before you pay
Read duration_seconds from the TTS result and round it up for the render. A render is billed per output minute, rounded up, so a 61-second result costs $0.20, not $0.10. A one-second trim from the script is worth a dime in that case.
Why this job is common now
More tools now narrate for you: Microsoft's MAI-Voice-2.1 lists voice cloning with consent guardrails and 23 languages (Microsoft AI, read 2026-10-04). That raises the number of old videos with a stale or off-brand voice. The swap is cheap, and each step returns its own job result, so you can stop after any step and review.
Two limits to remember. Detach takes sources up to 1,800 seconds and gives at most 900 seconds out, so a long video needs a range. And a source without an audio track fails with detach_source_has_no_audio. The jobs docs show how to poll each step.
Sources
Related posts
More in Media tools
- Restyle one captioned clip into three looks with source_caption_id
Pass source_caption_id to video-captions to re-burn the same clip in a new style without a second speech-to-text run. A restyle is still a billed render.
- Roku ad frame rates: only three allowed, probe first
Roku ads reportedly accept 23.98, 25 or 29.97 fps. Probe your clip before upload with Sume's video inspect, and cut it to length if it is outside 6 to 92 s.
- Samsung Ads mobile video, 10 MB recommended: bitrate for 15 and 30 s
Samsung recommends 10 MB for mobile and tablet video, 30 MB for desktop. That caps a 30-second spot near 2.7 Mbps. The math, plus a Sume render and size check.
- Seedance 2.5 secondary edit vs trim-and-regenerate on Sume
BytePlus describes timestamp-level edits to Seedance 2.5 clips. Sume has no such edit field, so here is the trim-and-regenerate route and its limits.
Written by Sume