Trim pauses from ten clips: inspect segments, then one timeline render
Rough-cut apps trim pauses in their own editor. For a batch by API, inspect each clip for sentence segments, keep the speech, and render one timeline.

To trim pauses from a batch of clips by API, inspect each clip with a transcript and sentence segmentation, keep the time ranges that contain speech, and send those ranges to one timeline render. Sume does not ship an automatic pause remover. It ships the pieces, an inspect that finds where sentences start and end, and a timeline that stitches ranges together, and you write the thirty lines between them.
In-app rough-cut tools now do this inside their own editors. If your clips live outside such an app, or you have ten of them rather than one, the scripted route is the one that scales. A related post covers the single-clip case: Instagram First Draft pause cutting outside the app.
Step 1: find the speech
Call POST /v1/video-inspect with transcribe true and a segmentation mode of sentence. Setting silence_split_seconds between 0.2 and 3 makes a gap of at least that long split a segment, so a pause becomes the space between two segments. The transcript adds $0.01 per audio minute to the inspect's own compute charge (read 2026-10-06 in the Sume docs), and a silent clip returns inspect_source_has_no_audio, so check probe.has_audio first with frames set to false (Sume docs: Video inspect, read 2026-10-06).
The clip must already be a media.sume.com artifact from this workspace. Import it with POST /v1/media-imports first, because the API does not fetch from the open internet.
Step 2: build one slot per segment
Each sentence segment has a start and an end in seconds. A timeline slot takes source_url, a start on the output spine, a duration and a source_in, the in-point into the file. So a segment from 3.4 to 7.9 becomes a slot with source_in 3.4 and duration 4.5, placed right after the previous slot. The first slot must start at 0 and later starts must increase.
A timeline does not play the audio of the clips. Its sound is its own audio spine, and audio.duration_seconds is required and sets the output length. To keep the speech you cut to, detach the audio of each clip once with POST /v1/audio-detach (flat $0.01 per job) and give audio.parts[] the same ranges as the video slots: url, source_in and duration for each. A parts list holds at most 20 entries, and the sum of the parts must not be shorter than duration_seconds, so set the duration to that sum. With more than 20 ranges across ten clips, join the audio first with a timeline-audio concat job (also at most 20 parts per job, $0.01 flat) or render clip by clip (audio detach, Timeline audio, Timeline 1.0, read 2026-10-06).
def ranges_for(segments, pad=0.08):
out = []
for seg in segments:
s = max(0.0, seg["start"] - pad)
d = round(seg["end"] + pad - s, 3)
if d >= 0.2:
out.append((round(s, 3), d))
return out
def build(video_url, audio_url, ranges):
video, parts, t = [], [], 0.0
for s, d in ranges:
video.append({"source_url": video_url, "start": round(t, 3),
"source_in": s, "duration": d})
parts.append({"url": audio_url, "source_in": s, "duration": d})
t += d
audio = {"parts": parts, "duration_seconds": round(t, 3)}
return {"audio": audio, "video": video}
segs = [{"start": 0.4, "end": 3.1}, {"start": 4.6, "end": 8.0}]
body = build("https://media.sume.com/artifacts/a/talk.mp4",
"https://media.sume.com/artifacts/a/talk.wav", ranges_for(segs))
print(len(body["video"]), "slots,", body["audio"]["duration_seconds"], "seconds")Step 3: render and check
Send the body to POST /v1/timeline-1.0/render with an Idempotency-Key, which is required. Rendering is $0.10 per ceiling output minute, and the plan endpoint is unbilled, so call the plan first to confirm the length and cost. A render allows up to 200 slots and 1,800 seconds.
Then watch the cut, at least once per batch. A pad of a tenth of a second keeps word starts from being clipped, and too little pad makes speech sound chopped.
| Step | Endpoint | Cost |
|---|---|---|
| Import clip | POST /v1/media-imports | See pricing |
| Find speech | POST /v1/video-inspect, transcribe | $0.01 per audio minute, plus inspect compute |
| Detach clip audio | POST /v1/audio-detach | $0.01 per job |
| Preview the cut | Timeline plan | Unbilled |
| Render the cut | POST /v1/timeline-1.0/render | $0.10 per output minute, rounded up |
Where this does not fit
If your clips are one continuous take with no spoken words, there are no sentence segments to key off, and the transcript step will report no audio or find nothing. For those, cut by timestamps you choose with video-trim. Music-only clips need a human decision about what to keep.
Sources
Related posts
More in Agents
- Weekly video: Agent Completion, Format or schedule on Sume?
A weekly video fits a Format or a schedule, not a fresh Agent Completion each time. How to choose among the three Sume surfaces for a recurring job.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
Written by Sume