Clicks or gaps when joining TTS MP3 clips: render WAV, join once

Joined MP3 voiceover clips can gap or click. Sume's Timeline audio docs explain why: MP3 adds priming padding at each edge. Keep WAV until the last step.

5 min readSume
All posts

If clips of a voiceover click or leave a tiny gap when you join them, the cause is usually the MP3 format and not the voice. The Timeline audio docs say WAV is the default for joins and splits because it is sample-exact (pcm_s16le), while MP3 adds priming padding at every edge it creates (Timeline audio, read 2026-10-06). So: render each clip as WAV, join them in one Timeline audio concat, and make the MP3 only once, at the very end, if you need one. Every re-encode of an MP3 adds another edge.

Faster speech models make the pieces smaller. Microsoft's MAI-Voice-2.1 Flash is listed with a vendor-claimed 150 ms end to end (October tracker, read 2026-10-06), and sentence-by-sentence pipelines are a common way to use that. The more clips you join, the more the format matters.

Why MP3 joins are not clean

An MP3 stream is made of frames, and encoders add a short run of padding at the start and end of a file so the decoder can settle. Cut an MP3 into pieces, or encode pieces one by one, and every piece has its own padding. Played back to back, the pieces can have a small silence or a discontinuity at each join. WAV with 16-bit PCM has no frames and no padding: the samples are the audio, so a cut at sample N is exactly at sample N.

Render WAV pieces, join, then decide the final format

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    out = json.load(u.urlopen(u.Request(url, data=data, headers=h)))
    return out.get("data", out)

LINES = ["First, the problem.", "Second, the fix.", "Third, the result."]
def done(job):
    sub = job
    while not job.get("terminal"):
        time.sleep(job.get("next_poll_after_seconds") or 2)
        job = call(sub["status_url"])
    res = call(sub["result_url"])
    return res.get("result", res)

jobs = [call("https://api.sume.com/v1/tts-router/generate", {
    "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
    "transcript": t, "output_format": {"container": "wav", "encoding": "pcm_s16le"}},
    f"wav-join-v1-{i}") for i, t in enumerate(LINES)]
parts = [{"url": next(a["url"] for a in done(j)["artifacts"] if a["type"] == "audio")} for j in jobs]
out = done(call("https://api.sume.com/v1/timeline-1.0/audio",
                {"operation": "concat", "parts": parts}, "wav-join-v1-concat"))
print(out.get("audio_url"), out.get("duration_seconds"))
for s in out.get("segments", []):
    print(s["index"], s["start"])

What you get

The concat output is WAV unless you ask otherwise, so the joined file is still sample-exact and you can hand it to the next tool, such as a lip sync model or a video timeline, without another lossy step. Only convert to MP3 when the destination needs it, and convert once. Output is limited to 1,800 seconds per job, and every part must be a media.sume.com URL, which TTS artifacts are.

Where to use each format in a voiceover pipeline (read 2026-10-06)
StepFormatReason
Render each sentenceWAV pcm_s16leNo padding, exact cuts
Join and trimWAVSample-exact edges
Lip sync or editor inputWAVNo extra lossy step
Final delivery to a playerMP3, onceSmall file; one set of edges

How to test a join

A practical check before you trust a join: play the joined file through the end of each clip boundary, listening for a click or a short hole, and compare it with the same text rendered as one single job. If the one-job version is clean and the joined one is not, look at the formats of the parts first. Mixed formats are the next suspect: the concat rejects parts whose channel layouts differ with audio_parts_channel_mismatch, so keep every part in the same container, rate and encoding.

Cost and related options

The cost is the same either way: TTS is $0.0475 per 1,000 characters regardless of container, and a join is a flat $0.01 (pricing, read 2026-10-06). WAV files are larger, so the only price of the cleaner route is storage and transfer for a short time. For cutting one long take into sentence clips instead of rendering pieces, see cut a voiceover into sentence clips. Job fields are in Jobs and results.

If you only need the join inside one video, skip the separate job. The Timeline audio docs point to Timeline 1.0 audio.parts[] for a join that lives inside a single render: up to 20 gapless slices, joined in the sample domain with no re-TTS, and no extra $0.01 job. Use the standalone concat when you want a reusable audio file, for example one narration that drives both a lip-sync clip and a video render, because it returns one audio_url plus segments[] with each part's start offset, which you can use to re-base your video slot starts.

Whichever route you pick, keep the parts in one channel layout. The concat refuses parts whose layouts differ with audio_parts_channel_mismatch, so do not mix a mono voice with a stereo music bed in the same call. Join the voice first, then mix the result with the bed in the render.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume