Gap or click when joining audio files: use wav, not mp3

Sume's Timeline audio joins up to 20 Sume-hosted parts with no silence at the seams. Its docs say to keep wav, because mp3 adds padding at every edge.

4 min readSume
All posts

To join audio files without a gap, join them as wav. Sume's Timeline audio (POST /v1/timeline-1.0/audio, operation concat) joins Sume-hosted audio in the sample domain, with no re-synthesis and no silence at the seams. Its schema says wav (pcm_s16le, the default) stays sample-exact, while mp3 is smaller but re-introduces priming padding at every edge.

The facts are from the Timeline audio request schema in the Sume API reference and the Timeline audio docs, read 2026-09-29.

What does the join do?

parts holds 1 to 20 Sume-hosted HTTPS audio URLs, joined in order. Each part may carry source_in and duration to cut before it joins. All parts must share one channel layout, or the request fails with audio_parts_channel_mismatch. The result is a durable media.sume.com file.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-join-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/example/line-1.wav" },
      { "url": "https://media.sume.com/artifacts/example/line-2.wav" }
    ],
    "output": { "format": "wav" }
  }'

Which output format should I ask for?

output.format is optional and defaults to wav. The schema gives two values.

From the Timeline audio request schema in the Sume API reference, read 2026-09-29.
`output.format`What the schema says
wav (default, pcm_s16le)Stays sample-exact, so the result can be joined again or used to drive Avatar 1.0 image-to-video
mp3Smaller, but re-introduces priming padding at every edge

Why does mp3 add a gap?

Sume's schema states the effect and not the cause: mp3 output re-introduces priming padding at every edge, so if you hear a pause at a seam in an mp3 join, that is the documented behavior to look at first. The wav default avoids the padding, so the result can be joined again or used to drive Avatar 1.0 image-to-video without picking up encoder delay.

Pick mp3 only for the last file you hand to a listener, and keep the working files wav. If you join twice, for example a voiceover line and then a music tail, the schema says the padding returns at every edge of an mp3, so join the wav files first and ask for mp3 only on the final join. Listen to the seams of your own files after the join; the docs make no numeric promise about how long the padding is.

What are the limits?

  • The join takes 1 to 20 parts per call, so a longer script needs a join of joins. Keep those intermediates wav too.
  • Only Sume-hosted URLs are accepted, per the schema. A part from anywhere else does not match the request schema, so join audio that already lives on media.sume.com in your workspace. Each part can still carry its own source_in and duration, so trimming happens in the same call.
  • If the joined audio is needed only inside one video, the docs say to use audio.parts[] on POST /v1/timeline-1.0/render and skip this job.
  • For layering music under a voice, see Mix voice with background music.

Sources

Related posts

More in Models

All Models posts

Written by Sume