Replace the voice in a finished video: detach, transcribe, TTS, render

Four Sume calls swap the narration in a finished video: detach the audio, transcribe it, voice your edited script, render. About 17 cents for one minute.

5 min readSume
All posts

To replace the narration in a finished video on Sume, chain four jobs: detach the audio from the video, transcribe it to get the script, voice your edited script with TTS, then render the video with the new audio as the spine. For a one-minute video with a 1,000-character script the rates add up to $0.1675: $0.01 detach, $0.01 transcription, $0.0475 TTS and $0.10 for one rendered minute (API reference).

Step by step

  • Detach: POST /v1/audio-detach takes a video_url on your workspace's media.sume.com. Import the file first if it lives elsewhere. Ask for channels: "mono" and sample_rate: 16000, which the detach docs call the speech-to-text shape.
  • Transcribe: POST /v1/stt-1.0/transcribe with the detached audio_url, up to 600 seconds per job. Edit the returned text.
  • Voice it: POST /v1/tts-1.0/generate with your edited text as a wav, so the render step receives a clean file.
  • Render: POST /v1/timeline-1.0/render with audio.url set to the TTS file, audio.duration_seconds set to its length, and the original clip as the single video[] slot (timeline docs).

The render step

The sound track of a Sume render comes from the audio you declare, plus an optional soundtrack, not from the clip. Still, listen to the output to confirm. If the new narration is longer than the clip, the render can pad or loop the video and says so in a warning, so keep the edited script close to the original length.

{
  "audio": {"url": "https://media.sume.com/artifacts/artf_new/voice.wav",
            "duration_seconds": 58},
  "video": [{"source_url": "https://media.sume.com/artifacts/artf_old/clip.mp4",
             "start": 0, "duration": 58, "fit": "cover"}]
}

Check the length before you pay

Read duration_seconds from the TTS result and round it up for the render. A render is billed per output minute, rounded up, so a 61-second result costs $0.20, not $0.10. A one-second trim from the script is worth a dime in that case.

Why this job is common now

More tools now narrate for you: Microsoft's MAI-Voice-2.1 lists voice cloning with consent guardrails and 23 languages (Microsoft AI, read 2026-10-04). That raises the number of old videos with a stale or off-brand voice. The swap is cheap, and each step returns its own job result, so you can stop after any step and review.

Two limits to remember. Detach takes sources up to 1,800 seconds and gives at most 900 seconds out, so a long video needs a range. And a source without an audio track fails with detach_source_has_no_audio. The jobs docs show how to poll each step.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume