34 pause cuts, 20 audio parts: split the spine in two levels

Timeline audio.parts holds 20 slices. For 34 pause cuts, build two concat files of 17, feed them as two parts, and re-base each video slot start.

5 min readSume
All posts

A Timeline render accepts at most 20 audio.parts[], so a talking clip cut at 34 pauses cannot slice its spine in one document. Detach the audio once, make two Timeline audio concat jobs of 17 parts each, and give the render those two files as its only two parts; the 34 video slots (limit 200) then use the concat segments[] offsets as their start values.

This matters when you automate what Instagram's First Draft does on a phone, cutting pauses in one tap (read 2026-10-09). A long take can have far more than 20 pauses.

The two-level plan

Step 1: POST /v1/audio-detach on the source clip, default wav, so the spine keeps the source rate and channels. Step 2: chunk the kept ranges into groups of at most 20 and post each group to POST /v1/timeline-1.0/audio with operation: "concat". Each part is {url, source_in, duration}, so one detached wav feeds every part. Step 3: render with audio.parts set to the concat outputs and audio.duration_seconds set to the summed duration.

Each concat result returns segments[] with index, start and duration_seconds. Add the running total of earlier concat files to each segment start; that is the on-spine start of the matching video slot.

import math

def plan(ranges, wav):
    """ranges: [(s, e)] on the detached wav. Chunks of 20 per concat job."""
    jobs = []
    for i in range(0, len(ranges), 20):
        parts = [{"url": wav, "source_in": s, "duration": round(e - s, 3)}
                 for s, e in ranges[i:i + 20]]
        jobs.append({"operation": "concat", "parts": parts})
    return jobs  # POST each to /v1/timeline-1.0/audio

def slots(ranges, src, concats):
    """concats: the result of each concat job (duration_seconds, segments)."""
    out, offset, i = [], 0.0, 0
    for c in concats:
        for seg in c["segments"]:
            s, _ = ranges[i]
            i += 1
            out.append({"source_url": src, "source_in": s,
                        "start": round(offset + seg["start"], 3),
                        "duration": seg["duration_seconds"]})
        offset += c["duration_seconds"]
    return out

How the job count grows

Audio cost is the detach ($0.01) plus $0.01 per concat job. The final render needs one audio.parts entry per concat file, and that list is capped at 20, so the 200-slot video limit binds first.

Audio jobs for N pause cuts (rates read 2026-10-09, arithmetic only)
CutsConcat jobsWithin 200 video slotsDetach + concat
200yes$0.01
342yes$0.03
1005yes$0.06
20010yes$0.11

Gotchas

  • For 20 cuts or fewer, skip concat: put the slices straight into audio.parts with source_in and duration, as in the Timeline-only version.
  • Every part must share a channel layout, or the concat fails with audio_parts_channel_mismatch. One detached wav avoids that.
  • Set audio.duration_seconds at or below the summed part length. If the parts hold less than the declared duration the render fails with audio_parts_shorter_than_duration, so round down, not up.
  • Do not use the 16 kHz transcript audio from video_inspect as the spine; the render flags a spine under 32 kHz as audio_spine_low_fidelity.
  • More than 12 slots renders chunked automatically; do not force render.strategy: "single".

Sources

Related posts

More in Developers

All Developers posts

Written by Sume