Replace the voice in a video you shot: detach, transcribe, re-voice

Swap the speech in a finished video on Sume: detach the audio, transcribe it, fix the script, speak it with TTS and render the picture with the new track.

4 min readSume
All posts

To replace the voice in a video you already shot, take the words out, correct them, speak them again and put the new track under the same picture. On Sume that is detach, transcribe, TTS and a Timeline 1.0 render. Use it for a flubbed line, an old voice you cannot use, or a clean read of a rough take. It does not move the lips, so use it where the speaker is off camera or small in frame.

Step 1: get the words

Import the video to your workspace, then call POST /v1/audio-detach. The default is a sample-exact wav, and the source can be up to 1800 seconds. Check probe.has_audio first, because a video with no audio track fails with detach_source_has_no_audio. Then send the wav to POST /v1/stt-1.0/transcribe, up to 600 seconds per file, with segmentation: {"mode": "sentence"}.

Step 2: correct and speak

Fix the transcript by hand. The segments tell you where each sentence sat in the original, which is the length you have to hit with the new voice. Speak each corrected sentence with POST /v1/tts-1.0/generate and compare the file's length with its segment. A line that runs long needs a shorter sentence.

Step 3: put it under the picture

Join the lines with timeline audio, then render with Timeline 1.0: audio.url is the joined file, audio.duration_seconds is its length, and the video[] slot is the original clip. Any original music or room sound goes with the old track, so add a bed back as a soundtrack if you want one. A render bills at $0.10 per output minute, rounded up.

The jobs and what they hold, read 2026-10-06:

Re-voice a finished video, read 2026-10-06
JobReadsReturns
audio-detachA media.sume.com videoA wav and its duration
stt-1.0A public HTTPS audio URLText, word timings, sentence segments
tts-1.0Your corrected sentenceAn audio file per line
timeline audio and renderJoined audio and a video slotOne MP4

When to leave the original voice

If the person is on camera and the mouth is clear, a different voice will look wrong, since nothing here moves the lips. Re-record the line instead, or cut to a shot where the mouth is not visible.

Also consider the sound that goes with the voice. Footsteps, a door and a room tone sit in the same track. After the swap the scene can sound empty, so add a quiet bed or a sound effect under the new voice.

Listen to the whole render before you publish. If a platform or a law asks you to label a synthetic voice, label it. Confirm the live rates in GET /v1/catalog.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume