Gaps between joined MP3 clips: priming padding and Sume's wav default
Joined MP3 clips can leave tiny gaps because each file carries priming padding. Sume Timeline audio defaults to sample-exact wav. When to use mp3.

The symptom
You render ten lines of speech, join them, and hear a small hitch at each seam. The cause is usually that the clips were MP3. The Timeline audio page says the mp3 output option re-adds priming padding at every edge, and recommends keeping wav when the file will be joined again or will drive lip-sync.
Compressed formats carry a bit of encoder padding at the start and often at the end. Join compressed files naively and the padding becomes silence between clips.
What Timeline audio does
Timeline audio's concat operation joins Sume-hosted audio in the sample domain. The docs say there is no re-synthesis and no silence at the seams. The default output is wav with pcm_s16le, described as sample-exact; mp3 is the smaller option.
| Output | Size | Seam behaviour | Use when |
|---|---|---|---|
| wav (default, pcm_s16le) | Larger | Sample-exact | Joining again, lip-sync, final master |
| mp3 | Smaller | Priming padding re-added at every edge | Final delivery of a finished file |
A pipeline that avoids gaps
Keep everything lossless until the last step.
- Ask Sume TTS for a wav
output_formatfor each line. - Import any outside files first with
POST /v1/media-imports, because Timeline audio accepts only this workspace's media.sume.com audio. - Concat with
operation: concatand the default wav output. - Encode to mp3 once, at the very end, or ship the wav.
Details that trip people up
Parts must share one channel layout, or the job returns audio_parts_channel_mismatch. A concat takes one to 20 parts, and produced audio is capped at 1,800 seconds. The price is a flat $0.01 per job whether you join two clips or twenty.
There is no GET resource for a timeline audio job. Poll the jobs envelope at /v1/jobs/:id/status and read /v1/jobs/:id/result, whose segments[] give you the offsets of each joined part.
How to confirm the problem is padding
Take two clips, join them as wav, and listen. Then encode each clip to mp3, join those, and listen at the same spot. If the first is clean and the second ticks, you have found it. Sume's own join does not insert silence at the seams, so the difference is the format.
The fix is order of operations: lossless join first, lossy encode last.
Takeaway
If your joins click, check the format first. Use wav through the whole chain and encode once. See wav vs mp3 for the general tradeoff.
Sources
Related posts
More in Developers
- Music API sync mode: wait_timeout_seconds 0-30, timed_out means poll
Sume music jobs accept mode sync with wait_timeout_seconds from 0 to 30. If sync.timed_out is true the job is still running; poll, never resubmit.
- Music API metadata field: tag tracks by campaign, not sent to provider
Sume's Music Router stores a metadata object on the job and does not send it to the provider. How to tag generations by campaign, brief or scene for audits.
- n8n browser OAuth2 webhook auth vs Sume HMAC-signed webhooks
n8n 2.42 adds a browser OAuth2 flow for User Auth webhooks. Sume's webhooks are server-to-server and HMAC-signed instead. A Python verifier that fails closed.
- n8n durable agent message queue: replays and Sume idempotency keys
n8n 2.42 lays a durable agent message queue foundation. A queue that can redeliver means a paid Sume call needs a stable idempotency_key per message.
Written by Sume