Join voiceover takes into one gapless wav: timeline audio concat
Timeline audio concat joins up to 20 hosted audio parts sample-exact into one reusable wav for $0.01 per job and returns segment offsets.

Send operation: "concat" and up to 20 parts[] to POST /v1/timeline-1.0/audio and you get one gapless file back at $0.01 per job. The join is in the sample domain, with no re-synthesis and no silence at the seams (Timeline audio).
Request
Each part is {url, source_in?, duration?}. The second part below skips its first 0.1 s and keeps 1.8 s.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: vo-concat-001" \
-d '{
"operation": "concat",
"parts": [
{"url": "https://media.sume.com/artifacts/artf_demo/line1.wav"},
{"url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8}
],
"output": {"format": "wav"}
}'What comes back
The result is kind: timeline_audio with one audio_url, duration_seconds and segments[] holding index, start and duration_seconds for each part. Use those offsets to re-base video[].start in your timeline.
Concat or audio.parts?
| Need | Use | Cost |
|---|---|---|
| Reusable merged file | timeline-audio concat | $0.01 per job |
| Join only inside one render | audio.parts[] on Timeline render | Included in the render |
| Slice one file into ranges | timeline-audio split | $0.01 per job |
Format notes
wav(pcm_s16le) is the default and sample-exact;mp3is smaller and adds priming padding at every edge, so keep wav if you will join again.- All parts must share a channel layout:
audio_parts_channel_mismatch. - The produced audio is at most 1800 s.
After the join
Pass the returned audio_url as audio.url on a Timeline 1.0 render, with audio.duration_seconds set to the returned duration_seconds. The same file also works as Avatar 1.0 image-to-video audio. If the render is the only place the voice is used, skip this job and send the takes as audio.parts[] instead.
Need the opposite? operation: "split" takes one top-level url and 1 to 20 ranges[], each {start, end?}, and returns a separate audio_url per segment. It is also $0.01 per job. If the source is a talking-head MP4, run audio detach once and split the audio it returns.
Sources
Related posts
More in Media tools
- LinkedIn Page video max ratio 2.4:1: crop a 32:9 recording
LinkedIn Pages accept video from 1:2.4 to 2.4:1 and up to 4096x2304. A 5120x1440 ultrawide capture fails both. Crop it with Sume video filter to fit.
- MAI-Voice 24 kHz 160 kbps MP3 vs Sume TTS 44.1 kHz 128 kbps: mixing
Microsoft's MAI-Voice example saves 24 kHz 160 kbps mono MP3; Sume TTS defaults to 44.1 kHz 128 kbps MP3 or WAV. What the numbers mean for mixing and file size.
- Music from a still: image_url conditioning, still $0.125
The Music Router accepts an optional public HTTPS image_url to condition the track on a still. Request example, keeping a score consistent, fixed price.
- Remove backgrounds from 500 product photos: RMBG API cost and code
Sume RMBG 1.0 removes a background for $0.0225 per image regardless of size. Batch cost for 50 to 5,000 photos, the request body and a Python loop.
Written by Sume