Concat 20 voice lines into one wav with timeline audio for 1 cent
Timeline audio concat joins up to 20 Sume-hosted audio parts into one gapless wav for a flat $0.01, and returns segment offsets to re-base your video slots.

Timeline audio operation: "concat" joins 1 to 20 ordered parts into one gapless file for a flat $0.01 per job. The join is sample-domain, with no re-synthesis and no silence at the seams. Twenty narrated lines cost one cent to merge, not 20.
What comes back
The result is kind: timeline_audio with one audio_url, duration_seconds and segments[] with index, start and duration_seconds. Those offsets re-base video[].start in a Timeline 1.0 render.
| Item | Value |
|---|---|
| Parts per job | 1 to 20 |
| Part fields | url, optional source_in, duration |
| Output format | wav (default, pcm_s16le) or mp3 |
| Output length | 1,800 s maximum |
| Price | $0.01 flat per job |
| Channel layouts | must match: audio_parts_channel_mismatch |
Request
Send parts and do not send a top-level url or ranges. All URLs must be audio already on media.sume.com. Import the clip first with POST /v1/media-imports so it sits on media.sume.com, send an Idempotency-Key header, then poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result. The API rejects off-host URLs at admit, so a bad URL fails before any work runs.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: concat-001" \
-d '{"operation":"concat","parts":[{"url":"https://media.sume.com/artifacts/artf_demo/line1.wav"},{"url":"https://media.sume.com/artifacts/artf_demo/line2.wav","source_in":0.1,"duration":1.8}]}'Concat or audio.parts
If the join is needed only inside one render, use audio.parts[] on Timeline 1.0, which takes up to 20 slices too. Use concat when you want a reusable file, for example as image-to-video audio for an avatar.
Keep wav
mp3 is smaller but adds priming padding at every edge. If the file will be joined again or drives lip-sync, keep wav. For more than 20 lines, concat in groups and concat the results: 2 groups of 20 plus one final join is 3 x $0.01 = $0.03 for 40 lines.
Using the offsets
Take the segments[] start values from the result and use them as video[].start in the render, so each picture changes where its line begins. With 20 lines of 3 seconds each, the offsets are 0, 3, 6, up to 57, and the file is 60 seconds, which renders as 1 minute at $0.10. The merge plus the render is $0.11 in total. Every media job follows the same lifecycle: submit with an Idempotency-Key, receive a job, poll GET /v1/jobs/:id/status until it is ready, then read GET /v1/jobs/:id/result. A retry with the same key does not queue a second job, so a network error during submit never doubles a charge.
Sources
Related posts
More in Media tools
- Conform a trim to 540x960 at 30 fps for TikTok non-Spark ads
TikTok's non-Spark ad spec sets 540x960 as the 9:16 minimum. One video-trim job with output 540x960 at 30 fps costs $0.02. Bitrate is not a field.
- Crop 1920x1080 to 3:4 with video filter: width 0.421875, x 0.2890625
A centered 3:4 crop of a 1920x1080 frame keeps 810x1080 pixels. In video-filter fractions that is width 0.421875, height 1, x 0.2890625. One job, $0.02.
- Crop 1920x1080 to 9:16 with an API: width 0.3164, x 0.3418
Center-crop a 16:9 clip to vertical with Sume video filter: the crop fractions, the arithmetic, and why 1280x720 falls under TikTok's 540x960 minimum.
- Crop a 1080x1920 clip to a 1080x1080 square for a TikTok 1:1 ad
TikTok's 1:1 ad needs 640x640 or more. A 9:16 source crops to a square with video filter height 0.5625 and y 0.21875, or y 0.1 to keep the top. $0.02.
Written by Sume