Join 20 voice lines: audio.parts in the render, or a $0.01 audio job
Timeline 1.0 takes up to 20 audio.parts in one render. A Timeline audio job joins up to 20 parts for $0.01 and returns offsets. Which to use; 45 lines.

To join 20 voice lines for one video, put them in audio.parts[] inside the Timeline 1.0 render. It takes up to 20 gapless slices and adds no separate fee. Use the Timeline audio job (POST /v1/timeline-1.0/audio, operation concat, $0.01 flat per job) when you need a reusable audio file or the concat offsets for placing video.
The two routes side by side
The Sume docs say it directly: if you need sliced voiceover only inside one render, use audio.parts[]. For a reusable merged file, use Timeline audio. Both join in the sample domain, with no re-synthesis and no silence at the seams.
| Item | audio.parts[] in the render | Timeline audio concat |
|---|---|---|
| Maximum parts | 20 | 20 (1 to 20, ordered) |
| Extra charge | none beyond the render, $0.10 per ceil(output minute) | $0.01 flat per job |
| Result | joined inside the render only | one durable media.sume.com file |
| Offsets | you compute them from known durations | segments[] with index, start, duration_seconds |
| Other uses | none | audio.url on a render; Avatar 1.0 image-to-video audio |
| Output format | the render's audio | wav (default, sample-exact) or mp3 |
Why the offsets matter
The concat result returns segments[]. They are the offsets that you use to re-base video[].start in the render. If your story has 20 lines of different lengths and one video slot per line, the offsets are what tell you where each slot begins. With audio.parts[] you can do the same sum yourself from the durations that you set with duration, but a concat job gives you the numbers from the actual joined file.
Use wav if the file will be joined again or drives lip-sync. The docs warn that mp3 adds priming padding at every edge.
When you have more than 20 lines
Both routes stop at 20 parts. A script with 45 lines needs ceil(45 / 20) = 3 groups. If you want the file route, run three concat jobs at $0.01 each, $0.03 in total, with groups of 20, 20 and 5 lines. Each result is a durable audio file. Then give the render audio.parts[] with three entries, or concat the three results once more for one file at another $0.01.
All parts must have the same channel layout, or the job fails with audio_parts_channel_mismatch. Every URL must already be a media.sume.com file of your workspace. Import first with POST /v1/media-imports.
import math
LINES = 45
PER_JOB = 20
FEE = 0.01
jobs = math.ceil(LINES / PER_JOB)
print("concat jobs:", jobs)
print("fee:", round(jobs * FEE, 2))
print("group sizes:", [min(PER_JOB, LINES - i * PER_JOB) for i in range(jobs)])What the audio route does not do
Neither route adds music, fades or ducking. If a bed must sit under the voice, use the soundtrack options on the render. The concat job also does not change loudness. If the 20 lines were recorded at different levels, the joined file keeps those differences, so normalise the source files before you join them.
A rule of thumb
- One render, one script of 20 lines or fewer:
audio.parts[]. - Several renders that share the same voiceover (for example a 9:16 and a 1:1 cut): one concat job, then
audio.urlon each render. - You need the start time of each line: the concat job's
segments[]. - More than 20 lines: group them first, then join the groups.
Sources
Related posts
More in Media tools
- Join voiceover takes into one gapless wav: timeline audio concat
Timeline audio concat joins up to 20 hosted audio parts sample-exact into one reusable wav for $0.01 per job and returns segment offsets.
- Just-listed banner above 25 property clips: 25 compose jobs, $0.50
Stack a just-listed banner image above a property walkthrough clip with Sume Timeline compose: $0.02 a shot, so 25 listings cost $0.50 before the final render.
- Keyframe or exact trim for a TikTok ad: which keeps 516 kbps?
Sume video-trim exact re-encodes with libx264; keyframe copies the stream. You cannot set bitrate, so measure size / duration against TikTok's 516 kbps floor.
- Kling 3 mute plus a Music 1.0 bed: $1.625 vs $2.10 with sound
A 10 s 9:16 Kling 3 clip is $1.40 mute or $2.10 with sound. Mute, one Music 1.0 bed ($0.125) and a $0.10 Timeline render total $1.625, 47.5 cents less.
Written by Sume