Join voice takes with no seam gaps: timeline audio concat, wav or mp3
Timeline audio concat joins up to 20 Sume-hosted takes in the sample domain, with no re-synthesis and no silence at the seams. Use wav if you will join again.

POST /v1/timeline-1.0/audio with operation: "concat" joins 1 to 20 Sume-hosted audio parts into one file in the sample domain. There is no re-synthesis and no silence at the seams. Output is wav by default; choose mp3 only for the final file, because mp3 adds priming padding at every edge. The price is a flat $0.01 per job.
Why join instead of re-synthesise
When a voiceover is made line by line, each line is its own file. Joining with a silence gap changes the pacing, and re-synthesising the whole script changes the voice. A sample-domain join keeps each take exactly as generated, and the result carries segments[] with the offset of every part.
Those offsets are what you need next. Use them to re-base video[].start on a Timeline 1.0 render, so each clip lands on its line.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timeline-audio-concat-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/line1.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8 }
]
}'Rules of the join
| Rule | Value |
|---|---|
| Parts | 1 to 20, ordered; each { url, source_in?, duration? } |
| Source | This workspace's media.sume.com audio |
| Channel layout | Same for all parts, else audio_parts_channel_mismatch |
Top-level url or ranges | Not allowed on concat: audio_concat_takes_no_url, audio_concat_takes_no_ranges |
| Output length | At most 1800 s |
| Price | $0.01 flat per job |
wav or mp3
output.format is wav (default, pcm_s16le, sample-exact) or mp3, which is smaller but adds priming padding again at every edge. Keep wav if you will join the file again, or if the file drives lip-sync. Convert to mp3 once, at the end.
Use the joined file as audio.url on a Timeline 1.0 render, or as audio input for an avatar.
Splitting is the same route
operation: "split" takes a top-level url and 1 to 20 ranges, each { start, end? }. Ranges may overlap, and a missing end runs to the end of the file. Each segment comes back with its own audio_url.
For many clips cut from a recording, detach the audio from the video once with POST /v1/audio-detach, then split that file here.
Join inside one render instead
If you need the join only inside a single render, put the slices in audio.parts[] of Timeline 1.0. That avoids a separate job. Use this route when you want a reusable file that you can listen to, re-use, or hand to another tool.
Related posts
More in Media tools
- Keep a whole 16:9 shot in 9:16: blur fill instead of cropping
Cropping loses 68% of a 16:9 frame. Timeline fit blur keeps the full shot centered over a blurred copy. How it works on Sume and when to prefer it.
- Kling Motion Control character_orientation: video or image?
Pick 'video' (the default) for complex motion up to 30 s; 'image' follows camera movement but fal documents a 10 s limit. Sume bills $0.1575 per second.
- Kling Motion Control prompt: appearance only, motion is the video
On Sume's Kling 3.0 Motion Control the prompt guides appearance only. Motion comes from motion_video_url, and the body rejects reference_image_urls and model.
- Lip-sync a saved avatar with H3 Max: avatar_handle, not image_url
Use a ready Sume avatar as the face for MiniMax H3 Max Lip Sync: send avatar_handle plus Sume-hosted audio, 5-14.8 seconds, and see the price at 768p.
Written by Sume