42 voice lines and the 20-part concat cap: four joins for 4 cents
Sume timeline audio joins at most 20 parts per job at $0.01 each. For 42 voice lines run three joins of 20, 20 and 2, then a final join, four jobs and $0.04.

Sume timeline audio concat accepts 1 to 20 parts per job, so 42 separate voice lines need more than one join. Join them as 20, 20 and 2 in three jobs, then join those three results in a fourth job. That is four jobs at a flat $0.01, or $0.04. The final file must stay within the 1,800-second limit on produced audio.
The limits and the price are from Sume's timeline audio docs and public catalog, read on 2026-10-09.
The plan
Concat is sample-domain: it joins Sume-hosted audio into one gapless file with no re-synthesis and no silence at the seams. Each output is a durable media.sume.com file with its own URL, so it is a valid input to another concat. All parts must share a channel layout, or the job fails with audio_parts_channel_mismatch.
| Step | Input | Output | Cost |
|---|---|---|---|
| Join A | lines 1 to 20 | audio A | $0.01 |
| Join B | lines 21 to 40 | audio B | $0.01 |
| Join C | lines 41 to 42 | audio C | $0.01 |
| Final join | A, B, C | one file | $0.01 |
| Total | 42 lines | 1 file | $0.04 |
What the request looks like
Each job is a POST to /v1/timeline-1.0/audio with operation concat and a parts array. Each part has a url and may have source_in and duration. All URLs must be your workspace's media.sume.com audio; an off-host URL is rejected at admission, so import any external files first. Idempotency-Key is required, and the default mode is async, so you get a 202 and poll the job unless you ask for sync, which waits up to 30 seconds.
The result is kind timeline_audio with one audio_url, a duration_seconds, and segments that give each part's start offset. Keep the segments from your joins if you need to line up video shots to the lines later.
When you do not need the intermediate files
If the lines will only be used inside one video render, you do not need to join them at all. Timeline 1.0 accepts audio.parts with up to 20 slices in a render. For more than 20, you still need an earlier join to get under the cap. And if the lines come from one TTS request with sentence segmentation, you already have the continuous take as audio_url, so joining slices is unnecessary.
The reusable merged file is useful when the same narration is used in several renders, since you pay the join once.
- Same channel layout on every part.
- Keep the order explicit in the parts array; there is no sort.
- Use a distinct Idempotency-Key per join, such as join-a, join-b, join-c, join-final.
Checking the result
After the final join, compare the duration_seconds of the output with the sum of your parts. The two should match within the usual rounding of your sources. A shortfall means a part was trimmed by a source_in or duration you did not intend. The segments array in each join result shows where each input landed, which is the quickest way to find the part that changed.
At $0.01 a job the check costs nothing to repeat, so rerun a join with a new Idempotency-Key if you need to change one input.
Sources
Related posts
More in Developers
- 61.4-second voice track on VEED Fabric: billed 62 s, $11.63 at 720p
VEED Fabric 1.0 on Sume bills ceil of audio seconds: 61.4 s is 62 s, 62 x $0.1875 = $11.625 at 720p, $6.20 at 480p. Includes a Python route picker.
- 7:5 and 5:7 images: crop from 3:2, 4:3, 3:4 or 2:3 on Sume
No Sume image row lists 7:5 or 5:7, but BFL's FLUX 3 Image does. How many pixels you lose cropping from 3:2, 4:3, 3:4 and 2:3, with the exact numbers.
- 8 avatar videos on a Free Sume workspace: 6 accepted, then 429
A Free Sume workspace runs 1 job, queues 5 and accepts 6. The 7th and 8th avatar video submits get 429 queue_full. A retry loop that waits and reuses the key.
- Eight Wan 3.0 clips on Free: six queue, two get queue_full
Free plan accepted capacity is 6 paid jobs. Submit eight 10-second Wan 3.0 720p jobs and two return 429 queue_full; the six accepted hold $7.50 of the $10.00.
Written by Sume