Join more than 20 voiceover lines: two-level timeline audio concat
Timeline audio concat takes 1 to 20 parts. A 120-line script needs six group joins and one final join, seven jobs, $0.07. A Python planner and offset math.

To join 120 per-sentence voiceover files you cannot send one request, because Sume's timeline audio concat takes 1 to 20 ordered parts. Do it in two levels: six concat jobs of 20 parts each, then one concat job whose six parts are the first-level outputs. That is seven jobs at $0.01 flat each, $0.07 in total, and the audio is still joined in the sample domain with no silence at the seams.
The part that trips people up is the offsets. A video that cuts on sentence starts needs absolute times, and each level reports only its own. This post shows the plan, the arithmetic and a runnable planner.
The limits that shape the plan
The wav default matters here. The docs note that mp3 re-adds priming padding at every edge, so an mp3 first-level output joined again would carry padding at each seam. Keep wav through both levels and convert once at the end, if you must.
| Rule | Value | Consequence |
|---|---|---|
| Parts per concat | 1 to 20 | 120 lines need two levels |
| Source of each part | This workspace's media.sume.com audio | Level-one outputs qualify as level-two inputs |
| Produced audio | Up to 1,800 seconds | The whole joined file must fit |
| Channel layout | All parts match (audio_parts_channel_mismatch) | Generate every line the same way |
| Output format | wav default, mp3 optional | Use wav for anything joined again |
| Price | $0.01 flat per job | 7 jobs, $0.07 |
| Poll | GET /v1/jobs/:id/status and /result | No resource URL of its own |
Offsets across two levels
Each concat result has segments[] with index, start and duration_seconds: the offsets of the parts inside that output. For two levels, a line's absolute start is its group's start in the final file plus its own start inside the group. The final job's segments[] gives the group starts. The first-level jobs give the line starts. Add them. Then use those absolute starts to re-base video slots in Timeline 1.0, as the docs describe for a single concat.
A planner you can run
This builds the first-level bodies and the final body with placeholder URLs, and then computes absolute line starts from line durations, which stand in for the values you will read back from each job's segments[].
import json
import math
N_LINES = 120
GROUP = 20
lines = [f"https://media.sume.com/artifacts/artf_demo/line{i:03d}.wav" for i in range(N_LINES)]
durations = [2.0 + (i % 5) * 0.4 for i in range(N_LINES)] # stand-in for measured seconds
groups = [lines[i:i + GROUP] for i in range(0, N_LINES, GROUP)]
level_one = [{"operation": "concat", "parts": [{"url": u} for u in g]} for g in groups]
print("level-one jobs:", len(level_one), "-> final job: 1 total $%.2f" % (0.01 * (len(level_one) + 1)))
# absolute line starts = group start + start inside group
group_len = [sum(durations[i * GROUP:(i + 1) * GROUP]) for i in range(len(groups))]
group_start = [sum(group_len[:i]) for i in range(len(groups))]
absolute = []
for gi in range(len(groups)):
inside = 0.0
for li in range(GROUP):
d = durations[gi * GROUP + li]
absolute.append(round(group_start[gi] + inside, 3))
inside += d
print("total seconds: %.1f (limit 1800)" % sum(durations))
print("line 0, 20, 119 start at:", absolute[0], absolute[20], absolute[119])
final_body = {"operation": "concat", "parts": [{"url": f"https://media.sume.com/artifacts/artf_group{i}/joined.wav"} for i in range(len(groups))]}
print(json.dumps(final_body)[:120], "...")
Order, names and idempotency
Give every request its own Idempotency-Key built from the group number, such as join-g3-v2, so a retry after a network drop does not create a second job, and bump the suffix when the content changes. Keep a manifest that records each group's job id, its parts in order, and its result URL; the final join is only as trustworthy as that list.
Finally, listen at the seams between groups, not just inside them. The join is sample-exact, so seams are clean, but a line that ended with a long breath or a clipped tail will be audible at any seam. Fix those at the line, not at the join.
Cheap ways to avoid the second level
If you only need the joined voiceover inside one render, skip the standalone job: Timeline 1.0 accepts audio.parts[] directly, and the docs say a join needed only inside one render belongs there. Use the two-level route when you need a reusable audio file, such as a podcast read, or when you want to hand the joined audio to another step like Avatar image-to-video.
Also consider retakes. If one of the 120 lines changes, rerun only that line and the one group that contains it, then rerun the final join: two jobs and a cent or two, not seven. Keep the first-level outputs and their segments[] around so a retake does not shift anything you have already timed.
Sources
Related posts
More in Developers
- Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.
- Count TTS characters like Sume: JS string length, emoji and Hangul
Sume TTS counts characters as JavaScript string length, so an emoji counts as 2. A short Python function counts UTF-16 units to predict the limit and cost.
- Cursor mcp.json ${env:NAME} for Sume's API key: no secret in the repo
Cursor's mcp.json interpolates ${env:NAME} in headers. Keep Sume's API key in an environment variable, send one credential, and know the fixed OAuth redirects.
- Cut a voiceover into sentence clips with TTS segmentation
Sume TTS returns gapless sentence segments, cutting 70 ms after each last word by default. Per-segment audio needs wav or raw; mp3 returns timings only.
Written by Sume