Clicks or gaps when joining TTS MP3 clips: render WAV, join once
Joined MP3 voiceover clips can gap or click. Sume's Timeline audio docs explain why: MP3 adds priming padding at each edge. Keep WAV until the last step.

If clips of a voiceover click or leave a tiny gap when you join them, the cause is usually the MP3 format and not the voice. The Timeline audio docs say WAV is the default for joins and splits because it is sample-exact (pcm_s16le), while MP3 adds priming padding at every edge it creates (Timeline audio, read 2026-10-06). So: render each clip as WAV, join them in one Timeline audio concat, and make the MP3 only once, at the very end, if you need one. Every re-encode of an MP3 adds another edge.
Faster speech models make the pieces smaller. Microsoft's MAI-Voice-2.1 Flash is listed with a vendor-claimed 150 ms end to end (October tracker, read 2026-10-06), and sentence-by-sentence pipelines are a common way to use that. The more clips you join, the more the format matters.
Why MP3 joins are not clean
An MP3 stream is made of frames, and encoders add a short run of padding at the start and end of a file so the decoder can settle. Cut an MP3 into pieces, or encode pieces one by one, and every piece has its own padding. Played back to back, the pieces can have a small silence or a discontinuity at each join. WAV with 16-bit PCM has no frames and no padding: the samples are the audio, so a cut at sample N is exactly at sample N.
Render WAV pieces, join, then decide the final format
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
out = json.load(u.urlopen(u.Request(url, data=data, headers=h)))
return out.get("data", out)
LINES = ["First, the problem.", "Second, the fix.", "Third, the result."]
def done(job):
sub = job
while not job.get("terminal"):
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(sub["status_url"])
res = call(sub["result_url"])
return res.get("result", res)
jobs = [call("https://api.sume.com/v1/tts-router/generate", {
"model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"transcript": t, "output_format": {"container": "wav", "encoding": "pcm_s16le"}},
f"wav-join-v1-{i}") for i, t in enumerate(LINES)]
parts = [{"url": next(a["url"] for a in done(j)["artifacts"] if a["type"] == "audio")} for j in jobs]
out = done(call("https://api.sume.com/v1/timeline-1.0/audio",
{"operation": "concat", "parts": parts}, "wav-join-v1-concat"))
print(out.get("audio_url"), out.get("duration_seconds"))
for s in out.get("segments", []):
print(s["index"], s["start"])What you get
The concat output is WAV unless you ask otherwise, so the joined file is still sample-exact and you can hand it to the next tool, such as a lip sync model or a video timeline, without another lossy step. Only convert to MP3 when the destination needs it, and convert once. Output is limited to 1,800 seconds per job, and every part must be a media.sume.com URL, which TTS artifacts are.
| Step | Format | Reason |
|---|---|---|
| Render each sentence | WAV pcm_s16le | No padding, exact cuts |
| Join and trim | WAV | Sample-exact edges |
| Lip sync or editor input | WAV | No extra lossy step |
| Final delivery to a player | MP3, once | Small file; one set of edges |
How to test a join
A practical check before you trust a join: play the joined file through the end of each clip boundary, listening for a click or a short hole, and compare it with the same text rendered as one single job. If the one-job version is clean and the joined one is not, look at the formats of the parts first. Mixed formats are the next suspect: the concat rejects parts whose channel layouts differ with audio_parts_channel_mismatch, so keep every part in the same container, rate and encoding.
Cost and related options
The cost is the same either way: TTS is $0.0475 per 1,000 characters regardless of container, and a join is a flat $0.01 (pricing, read 2026-10-06). WAV files are larger, so the only price of the cleaner route is storage and transfer for a short time. For cutting one long take into sentence clips instead of rendering pieces, see cut a voiceover into sentence clips. Job fields are in Jobs and results.
If you only need the join inside one video, skip the separate job. The Timeline audio docs point to Timeline 1.0 audio.parts[] for a join that lives inside a single render: up to 20 gapless slices, joined in the sample domain with no re-TTS, and no extra $0.01 job. Use the standalone concat when you want a reusable audio file, for example one narration that drives both a lip-sync clip and a video render, because it returns one audio_url plus segments[] with each part's start offset, which you can use to re-base your video slot starts.
Whichever route you pick, keep the parts in one channel layout. The concat refuses parts whose layouts differ with audio_parts_channel_mismatch, so do not mix a mono voice with a stereo music bed in the same call. Join the voice first, then mix the result with the bed in the render.
Sources
Related posts
More in Developers
- Postgres SKIP LOCKED poll table for AI video jobs with next_poll_at
Track Sume video jobs in one Postgres table with next_poll_at, claim due rows with FOR UPDATE SKIP LOCKED and reschedule from next_poll_after_seconds.
- pytest: fail CI when a configured image model id isn't in Sume's list
A pytest that reads your image model ids and checks each against GET /v1/images/models, so a retired id such as gpt-image-1 fails CI before it fails a customer.
- Python asyncio semaphore: submit transcription jobs with a cap
Run Sume STT submits through asyncio.Semaphore and asyncio.to_thread, honor retry-after on 429, and keep one Idempotency-Key per clip. Tested code.
- Python exceptions for Sume errors: retry on the class, not the status
Map the Sume error envelope code to two exception classes, Retryable and Fatal, so one except clause drives retries. Covers 409, 429, 402 and 503.
Written by Sume