Retake one sentence of a voiceover and splice it in with concat parts

Re-record one sentence, not the whole TTS take: use wav sentence segments and one timeline audio concat job with source_in. Python for the parts list.

5 min readSume
All posts

Ask for wav output, word timestamps and sentence segmentation on the first take, so you know where each sentence starts and ends. To fix sentence k, synthesize that sentence alone, then run one timeline audio concat job with three parts: the original up to the sentence, the retake, and the original from the sentence's end. The join is sample-domain and gapless.

The retake will rarely match the old sentence's length, so everything after it moves. Re-base any captions or video starts from the concat result.

What do I ask for on the first take?

Send timestamps.words true, segmentation.mode sentence and output_format with container wav. The segments come back gapless (each end equals the next start), with a 70 millisecond lead after each sentence's last word by default, tunable with boundary_lead_ms from 0 to 500. With emit_audio on and wav or raw output, each segment also has its own audio_url. With mp3 you get the timings without per-segment audio, and the docs say to keep wav for audio that will be joined again.

How do I build the three parts?

Concat takes 1 to 20 ordered parts, each with a url and optional source_in and duration. All URLs must be media.sume.com audio from your workspace, which the TTS artifacts are. Using source_in on the full take means you need three parts however long the script is, so you stay under the 20-part ceiling.

The function below builds the list for any sentence index. It runs as written.

def splice_parts(full_url, retake_url, segments, k, retake_seconds):
    """Parts for a concat job that swaps sentence k for a retake."""
    start, end = segments[k]["start"], segments[k]["end"]
    parts = []
    if start > 0:
        parts.append({"url": full_url, "source_in": 0, "duration": start})
    parts.append({"url": retake_url, "duration": retake_seconds})
    parts.append({"url": full_url, "source_in": end})
    return parts

segments = [{"start": 0.0, "end": 3.1}, {"start": 3.1, "end": 6.8}, {"start": 6.8, "end": 10.2}]
for p in splice_parts("https://media.sume.com/artifacts/artf_demo/vo.wav",
                      "https://media.sume.com/artifacts/artf_demo/retake.wav",
                      segments, 1, 3.4):
    print(p)

What do I get back and what does it cost?

The concat job returns one audio_url, a duration_seconds and segments[] with the offsets of each part. Use those offsets to re-base caption cues and Timeline video starts. The produced audio can run up to 1,800 seconds, and the job is $0.01 flat per the timeline audio docs, plus the retake's own TTS charge.

Use the same voice, language and speed settings for the retake, and read the previous job's request to copy them.

The three parts for sentence 2 of 3 (output of the sketch, read 2026-10-02).
Parturlsource_induration
1full take03.1
2retakenone3.4
3full take6.8rest of file

What can still go wrong?

A retake can differ in pace, volume or breath from its neighbours, so listen to both joins. Tone is the common miss. If it sounds wrong, retake again with the same settings, or retake the sentence before and after it as one unit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume