Trim pauses from ten clips: inspect segments, then one timeline render

Rough-cut apps trim pauses in their own editor. For a batch by API, inspect each clip for sentence segments, keep the speech, and render one timeline.

5 min readSume
All posts

To trim pauses from a batch of clips by API, inspect each clip with a transcript and sentence segmentation, keep the time ranges that contain speech, and send those ranges to one timeline render. Sume does not ship an automatic pause remover. It ships the pieces, an inspect that finds where sentences start and end, and a timeline that stitches ranges together, and you write the thirty lines between them.

In-app rough-cut tools now do this inside their own editors. If your clips live outside such an app, or you have ten of them rather than one, the scripted route is the one that scales. A related post covers the single-clip case: Instagram First Draft pause cutting outside the app.

Step 1: find the speech

Call POST /v1/video-inspect with transcribe true and a segmentation mode of sentence. Setting silence_split_seconds between 0.2 and 3 makes a gap of at least that long split a segment, so a pause becomes the space between two segments. The transcript adds $0.01 per audio minute to the inspect's own compute charge (read 2026-10-06 in the Sume docs), and a silent clip returns inspect_source_has_no_audio, so check probe.has_audio first with frames set to false (Sume docs: Video inspect, read 2026-10-06).

The clip must already be a media.sume.com artifact from this workspace. Import it with POST /v1/media-imports first, because the API does not fetch from the open internet.

Step 2: build one slot per segment

Each sentence segment has a start and an end in seconds. A timeline slot takes source_url, a start on the output spine, a duration and a source_in, the in-point into the file. So a segment from 3.4 to 7.9 becomes a slot with source_in 3.4 and duration 4.5, placed right after the previous slot. The first slot must start at 0 and later starts must increase.

A timeline does not play the audio of the clips. Its sound is its own audio spine, and audio.duration_seconds is required and sets the output length. To keep the speech you cut to, detach the audio of each clip once with POST /v1/audio-detach (flat $0.01 per job) and give audio.parts[] the same ranges as the video slots: url, source_in and duration for each. A parts list holds at most 20 entries, and the sum of the parts must not be shorter than duration_seconds, so set the duration to that sum. With more than 20 ranges across ten clips, join the audio first with a timeline-audio concat job (also at most 20 parts per job, $0.01 flat) or render clip by clip (audio detach, Timeline audio, Timeline 1.0, read 2026-10-06).

def ranges_for(segments, pad=0.08):
    out = []
    for seg in segments:
        s = max(0.0, seg["start"] - pad)
        d = round(seg["end"] + pad - s, 3)
        if d >= 0.2:
            out.append((round(s, 3), d))
    return out


def build(video_url, audio_url, ranges):
    video, parts, t = [], [], 0.0
    for s, d in ranges:
        video.append({"source_url": video_url, "start": round(t, 3),
                      "source_in": s, "duration": d})
        parts.append({"url": audio_url, "source_in": s, "duration": d})
        t += d
    audio = {"parts": parts, "duration_seconds": round(t, 3)}
    return {"audio": audio, "video": video}


segs = [{"start": 0.4, "end": 3.1}, {"start": 4.6, "end": 8.0}]
body = build("https://media.sume.com/artifacts/a/talk.mp4",
             "https://media.sume.com/artifacts/a/talk.wav", ranges_for(segs))
print(len(body["video"]), "slots,", body["audio"]["duration_seconds"], "seconds")

Step 3: render and check

Send the body to POST /v1/timeline-1.0/render with an Idempotency-Key, which is required. Rendering is $0.10 per ceiling output minute, and the plan endpoint is unbilled, so call the plan first to confirm the length and cost. A render allows up to 200 slots and 1,800 seconds.

Then watch the cut, at least once per batch. A pad of a tenth of a second keeps word starts from being clipped, and too little pad makes speech sound chopped.

Pieces of a scripted pause trim and their cost, read 2026-10-06 against Sume docs
StepEndpointCost
Import clipPOST /v1/media-importsSee pricing
Find speechPOST /v1/video-inspect, transcribe$0.01 per audio minute, plus inspect compute
Detach clip audioPOST /v1/audio-detach$0.01 per job
Preview the cutTimeline planUnbilled
Render the cutPOST /v1/timeline-1.0/render$0.10 per output minute, rounded up

Where this does not fit

If your clips are one continuous take with no spoken words, there are no sentence segments to key off, and the transcript step will report no audio or find nothing. For those, cut by timestamps you choose with video-trim. Music-only clips need a human decision about what to keep.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume