34 pause cuts, 20 audio parts: split the spine in two levels
Timeline audio.parts holds 20 slices. For 34 pause cuts, build two concat files of 17, feed them as two parts, and re-base each video slot start.

A Timeline render accepts at most 20 audio.parts[], so a talking clip cut at 34 pauses cannot slice its spine in one document. Detach the audio once, make two Timeline audio concat jobs of 17 parts each, and give the render those two files as its only two parts; the 34 video slots (limit 200) then use the concat segments[] offsets as their start values.
This matters when you automate what Instagram's First Draft does on a phone, cutting pauses in one tap (read 2026-10-09). A long take can have far more than 20 pauses.
The two-level plan
Step 1: POST /v1/audio-detach on the source clip, default wav, so the spine keeps the source rate and channels. Step 2: chunk the kept ranges into groups of at most 20 and post each group to POST /v1/timeline-1.0/audio with operation: "concat". Each part is {url, source_in, duration}, so one detached wav feeds every part. Step 3: render with audio.parts set to the concat outputs and audio.duration_seconds set to the summed duration.
Each concat result returns segments[] with index, start and duration_seconds. Add the running total of earlier concat files to each segment start; that is the on-spine start of the matching video slot.
import math
def plan(ranges, wav):
"""ranges: [(s, e)] on the detached wav. Chunks of 20 per concat job."""
jobs = []
for i in range(0, len(ranges), 20):
parts = [{"url": wav, "source_in": s, "duration": round(e - s, 3)}
for s, e in ranges[i:i + 20]]
jobs.append({"operation": "concat", "parts": parts})
return jobs # POST each to /v1/timeline-1.0/audio
def slots(ranges, src, concats):
"""concats: the result of each concat job (duration_seconds, segments)."""
out, offset, i = [], 0.0, 0
for c in concats:
for seg in c["segments"]:
s, _ = ranges[i]
i += 1
out.append({"source_url": src, "source_in": s,
"start": round(offset + seg["start"], 3),
"duration": seg["duration_seconds"]})
offset += c["duration_seconds"]
return outHow the job count grows
Audio cost is the detach ($0.01) plus $0.01 per concat job. The final render needs one audio.parts entry per concat file, and that list is capped at 20, so the 200-slot video limit binds first.
| Cuts | Concat jobs | Within 200 video slots | Detach + concat |
|---|---|---|---|
| 20 | 0 | yes | $0.01 |
| 34 | 2 | yes | $0.03 |
| 100 | 5 | yes | $0.06 |
| 200 | 10 | yes | $0.11 |
Gotchas
- For 20 cuts or fewer, skip concat: put the slices straight into
audio.partswithsource_inandduration, as in the Timeline-only version. - Every part must share a channel layout, or the concat fails with
audio_parts_channel_mismatch. One detached wav avoids that. - Set
audio.duration_secondsat or below the summed part length. If the parts hold less than the declared duration the render fails withaudio_parts_shorter_than_duration, so round down, not up. - Do not use the 16 kHz transcript audio from
video_inspectas the spine; the render flags a spine under 32 kHz asaudio_spine_low_fidelity. - More than 12 slots renders chunked automatically; do not force
render.strategy: "single".
Sources
Related posts
More in Developers
- 37 video jobs submitted at once: which Sume plan accepts them all
Free accepts 6 of 37 video jobs, Pro 24, Startup all 37 with 11 spare slots, Scale all 37. Concurrency, queue and accepted-capacity table from the Sume docs.
- 4 of Sume's 17 aspect ratios exceed 3:1, GPT Image 2.5's size cap
Sume normalizes 17 aspect ratios; 1:4, 4:1, 1:8 and 8:1 are wider than the 3:1 limit on GPT Image 2.5 custom sizes. A check script and the 2 other rules.
- 42 voice lines and the 20-part concat cap: four joins for 4 cents
Sume timeline audio joins at most 20 parts per job at $0.01 each. For 42 voice lines run three joins of 20, 20 and 2, then a final join, four jobs and $0.04.
- 61.4-second voice track on VEED Fabric: billed 62 s, $11.63 at 720p
VEED Fabric 1.0 on Sume bills ceil of audio seconds: 61.4 s is 62 s, 62 x $0.1875 = $11.625 at 720p, $6.20 at 480p. Includes a Python route picker.
Written by Sume