Join dubbed voiceover lines into one gapless file for a cent on Sume
Sume timeline audio concatenates up to 20 audio parts or splits one file into up to 20 ranges, no re-synthesis and no silence at the seams, at $0.01 a job.

Concatenate them. Sume timeline audio joins 1 to 20 ordered parts into one file for $0.01 flat. The join happens in the sample domain, so there is no re-synthesis and no silence at the seams, and you get back a durable media.sume.com URL plus the offsets where each part landed.
Concat, step by step
- operation: concat with parts[]. Each part is a url, with optional source_in and duration to take a slice.
- All URLs must already be this workspace's media.sume.com audio. Import first.
- Parts must share one channel layout, or you get audio_parts_channel_mismatch.
- Default mode is async. mode sync waits up to 30 seconds for a 200.
- Output is up to 1800 seconds.
Use the segment offsets
The result carries segments[] with index, start and duration_seconds for each part. Those are the offsets to re-base your video start times against, so a dubbed line lands under the shot it belongs to. If a join is needed only inside one render, put it in the render's audio.parts instead and skip this job.
Split the other way
To cut many ranges out of a talking-head MP4, run audio detach once ($0.01), then split here with up to 20 ranges, $0.01. Ranges may overlap, and an omitted end means the rest of the file. A 20-line script costs $0.02 to detach and split.
| Operation | Input | Limit | Cost |
|---|---|---|---|
| concat | parts[] | 1 to 20 parts | $0.01 |
| split | url plus ranges[] | 1 to 20 ranges | $0.01 |
| wav (default) | pcm_s16le | Sample exact | Same |
| mp3 | Smaller file | Re-adds padding at each edge | Same |
Keep wav when you join again
The docs advise wav when the file will be joined again or drives lip-sync; mp3 re-adds priming padding at every edge.
Concat request
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: join-1" \
-d '{"operation":"concat","parts":[{"url":"https://media.sume.com/artifacts/artf_demo/line1.wav"},{"url":"https://media.sume.com/artifacts/artf_demo/line2.wav"}]}'Sources
Related posts
More in Media tools
- Keep a vertical video out of Shorts: render a 16:9 blur-fill cut
YouTube's help names a wider ratio such as 16:9 as the way to avoid Shorts. Put a vertical clip in a 1920x1080 frame with Sume's blur fit, $0.10 a minute.
- Put music under Gemini Omni's native audio: detach, spine, duck
Omni clips always carry native audio on Sume. Detach it to a wav, use it as the Timeline spine, add a music soundtrack with duck_db. Calls and limits.
- Burning Korean captions: why a Latin style returns a 400 on Sume
Sume's slam, punch and tiktok-green styles have no Hangul glyphs, so Korean copy on them fails with a 400. Which style to pick and what a $0.20 caption costs.
- Lip Sync or Native Audio for a Talking Shot: Which Sume Route
Sume's docs say video models do not lip-sync to TTS. Compare Fabric, H3 Max lip sync, Avatar Video and Kling native lip sync for a speaking shot.
Written by Sume