Replace the voice in a video you shot: detach, transcribe, re-voice
Swap the speech in a finished video on Sume: detach the audio, transcribe it, fix the script, speak it with TTS and render the picture with the new track.

To replace the voice in a video you already shot, take the words out, correct them, speak them again and put the new track under the same picture. On Sume that is detach, transcribe, TTS and a Timeline 1.0 render. Use it for a flubbed line, an old voice you cannot use, or a clean read of a rough take. It does not move the lips, so use it where the speaker is off camera or small in frame.
Step 1: get the words
Import the video to your workspace, then call POST /v1/audio-detach. The default is a sample-exact wav, and the source can be up to 1800 seconds. Check probe.has_audio first, because a video with no audio track fails with detach_source_has_no_audio. Then send the wav to POST /v1/stt-1.0/transcribe, up to 600 seconds per file, with segmentation: {"mode": "sentence"}.
Step 2: correct and speak
Fix the transcript by hand. The segments tell you where each sentence sat in the original, which is the length you have to hit with the new voice. Speak each corrected sentence with POST /v1/tts-1.0/generate and compare the file's length with its segment. A line that runs long needs a shorter sentence.
Step 3: put it under the picture
Join the lines with timeline audio, then render with Timeline 1.0: audio.url is the joined file, audio.duration_seconds is its length, and the video[] slot is the original clip. Any original music or room sound goes with the old track, so add a bed back as a soundtrack if you want one. A render bills at $0.10 per output minute, rounded up.
The jobs and what they hold, read 2026-10-06:
| Job | Reads | Returns |
|---|---|---|
| audio-detach | A media.sume.com video | A wav and its duration |
| stt-1.0 | A public HTTPS audio URL | Text, word timings, sentence segments |
| tts-1.0 | Your corrected sentence | An audio file per line |
| timeline audio and render | Joined audio and a video slot | One MP4 |
When to leave the original voice
If the person is on camera and the mouth is clear, a different voice will look wrong, since nothing here moves the lips. Re-record the line instead, or cut to a shot where the mouth is not visible.
Also consider the sound that goes with the voice. Footsteps, a door and a room tone sit in the same track. After the swap the scene can sound empty, so add a quiet bed or a sound effect under the new voice.
Listen to the whole render before you publish. If a platform or a law asks you to label a synthetic voice, label it. Confirm the live rates in GET /v1/catalog.
Sources
Related posts
More in Use cases
- AI avatar follow-up videos for stalled B2B deals: what they can do
Tavus cites Gartner on stalled B2B buying. A recorded avatar follow-up can restate one point, not answer questions. A 3-scene request and the cost by tier.
- Baby shower invitation art with an API: watercolor, details in code
Make a soft watercolor invitation border on Sume with an empty center, then add the date, place and host in code so every detail is exact and easy to update.
- Black Friday grocery and toy promo videos with AI: where to start
Adobe's forecast has grocery up 10.3% and toys up 9.6%, the fastest of its categories. Which Sume Formats fit a food or toy promo, and what to check.
- Black Friday video ads for phones: which Sume Formats are vertical
Adobe expects 57.4% of holiday online sales on mobile. Which Sume catalog Formats are described as vertical, and how to ask for vertical in the rest.
Written by Sume