Join more than 20 voiceover lines: two-level timeline audio concat

Timeline audio concat takes 1 to 20 parts. A 120-line script needs six group joins and one final join, seven jobs, $0.07. A Python planner and offset math.

6 min readSume
All posts

To join 120 per-sentence voiceover files you cannot send one request, because Sume's timeline audio concat takes 1 to 20 ordered parts. Do it in two levels: six concat jobs of 20 parts each, then one concat job whose six parts are the first-level outputs. That is seven jobs at $0.01 flat each, $0.07 in total, and the audio is still joined in the sample domain with no silence at the seams.

The part that trips people up is the offsets. A video that cuts on sentence starts needs absolute times, and each level reports only its own. This post shows the plan, the arithmetic and a runnable planner.

The limits that shape the plan

The wav default matters here. The docs note that mp3 re-adds priming padding at every edge, so an mp3 first-level output joined again would carry padding at each seam. Keep wav through both levels and convert once at the end, if you must.

From Sume's Timeline audio page, read 2026-10-03.
RuleValueConsequence
Parts per concat1 to 20120 lines need two levels
Source of each partThis workspace's media.sume.com audioLevel-one outputs qualify as level-two inputs
Produced audioUp to 1,800 secondsThe whole joined file must fit
Channel layoutAll parts match (audio_parts_channel_mismatch)Generate every line the same way
Output formatwav default, mp3 optionalUse wav for anything joined again
Price$0.01 flat per job7 jobs, $0.07
PollGET /v1/jobs/:id/status and /resultNo resource URL of its own

Offsets across two levels

Each concat result has segments[] with index, start and duration_seconds: the offsets of the parts inside that output. For two levels, a line's absolute start is its group's start in the final file plus its own start inside the group. The final job's segments[] gives the group starts. The first-level jobs give the line starts. Add them. Then use those absolute starts to re-base video slots in Timeline 1.0, as the docs describe for a single concat.

A planner you can run

This builds the first-level bodies and the final body with placeholder URLs, and then computes absolute line starts from line durations, which stand in for the values you will read back from each job's segments[].

import json
import math

N_LINES = 120
GROUP = 20
lines = [f"https://media.sume.com/artifacts/artf_demo/line{i:03d}.wav" for i in range(N_LINES)]
durations = [2.0 + (i % 5) * 0.4 for i in range(N_LINES)]  # stand-in for measured seconds

groups = [lines[i:i + GROUP] for i in range(0, N_LINES, GROUP)]
level_one = [{"operation": "concat", "parts": [{"url": u} for u in g]} for g in groups]
print("level-one jobs:", len(level_one), "-> final job: 1  total $%.2f" % (0.01 * (len(level_one) + 1)))

# absolute line starts = group start + start inside group
group_len = [sum(durations[i * GROUP:(i + 1) * GROUP]) for i in range(len(groups))]
group_start = [sum(group_len[:i]) for i in range(len(groups))]
absolute = []
for gi in range(len(groups)):
    inside = 0.0
    for li in range(GROUP):
        d = durations[gi * GROUP + li]
        absolute.append(round(group_start[gi] + inside, 3))
        inside += d

print("total seconds: %.1f (limit 1800)" % sum(durations))
print("line 0, 20, 119 start at:", absolute[0], absolute[20], absolute[119])
final_body = {"operation": "concat", "parts": [{"url": f"https://media.sume.com/artifacts/artf_group{i}/joined.wav"} for i in range(len(groups))]}
print(json.dumps(final_body)[:120], "...")

Order, names and idempotency

Give every request its own Idempotency-Key built from the group number, such as join-g3-v2, so a retry after a network drop does not create a second job, and bump the suffix when the content changes. Keep a manifest that records each group's job id, its parts in order, and its result URL; the final join is only as trustworthy as that list.

Finally, listen at the seams between groups, not just inside them. The join is sample-exact, so seams are clean, but a line that ended with a long breath or a clipped tail will be audible at any seam. Fix those at the line, not at the join.

Cheap ways to avoid the second level

If you only need the joined voiceover inside one render, skip the standalone job: Timeline 1.0 accepts audio.parts[] directly, and the docs say a join needed only inside one render belongs there. Use the two-level route when you need a reusable audio file, such as a podcast read, or when you want to hand the joined audio to another step like Avatar image-to-video.

Also consider retakes. If one of the 120 lines changes, rerun only that line and the one group that contains it, then rerun the final join: two jobs and a cent or two, not seven. Keep the first-level outputs and their segments[] around so a retake does not shift anything you have already timed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume