Join voice takes with no seam gaps: timeline audio concat, wav or mp3

Timeline audio concat joins up to 20 Sume-hosted takes in the sample domain, with no re-synthesis and no silence at the seams. Use wav if you will join again.

5 min readSume
All posts

POST /v1/timeline-1.0/audio with operation: "concat" joins 1 to 20 Sume-hosted audio parts into one file in the sample domain. There is no re-synthesis and no silence at the seams. Output is wav by default; choose mp3 only for the final file, because mp3 adds priming padding at every edge. The price is a flat $0.01 per job.

Why join instead of re-synthesise

When a voiceover is made line by line, each line is its own file. Joining with a silence gap changes the pacing, and re-synthesising the whole script changes the voice. A sample-domain join keeps each take exactly as generated, and the result carries segments[] with the offset of every part.

Those offsets are what you need next. Use them to re-base video[].start on a Timeline 1.0 render, so each clip lands on its line.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: timeline-audio-concat-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/line1.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8 }
    ]
  }'

Rules of the join

From docs.sume.com/models/timeline-audio, read 2026-10-05
RuleValue
Parts1 to 20, ordered; each { url, source_in?, duration? }
SourceThis workspace's media.sume.com audio
Channel layoutSame for all parts, else audio_parts_channel_mismatch
Top-level url or rangesNot allowed on concat: audio_concat_takes_no_url, audio_concat_takes_no_ranges
Output lengthAt most 1800 s
Price$0.01 flat per job

wav or mp3

output.format is wav (default, pcm_s16le, sample-exact) or mp3, which is smaller but adds priming padding again at every edge. Keep wav if you will join the file again, or if the file drives lip-sync. Convert to mp3 once, at the end.

Use the joined file as audio.url on a Timeline 1.0 render, or as audio input for an avatar.

Splitting is the same route

operation: "split" takes a top-level url and 1 to 20 ranges, each { start, end? }. Ranges may overlap, and a missing end runs to the end of the file. Each segment comes back with its own audio_url.

For many clips cut from a recording, detach the audio from the video once with POST /v1/audio-detach, then split that file here.

Join inside one render instead

If you need the join only inside a single render, put the slices in audio.parts[] of Timeline 1.0. That avoids a separate job. Use this route when you want a reusable file that you can listen to, re-use, or hand to another tool.

Related posts

More in Media tools

All Media tools posts

Written by Sume