Event recap: trim dead pauses from a talk with inspect and trim

YouTube's coming editing assistant promises to trim pauses. Until then, find silent gaps in a talk from transcript segments and cut keeper ranges with Sume.

5 min readSume
All posts

To trim dead pauses out of an event talk, transcribe it with Sume video inspect using sentence segments, find the gaps between segments that run past a threshold, then cut the keeper ranges with video trim and join them in a Timeline render. YouTube's blog says a conversational editing tool for Shorts and the YouTube Create app is coming and will help trim pauses and reorder for pacing, but gives no date.

The YouTube line comes from its Made on YouTube 2026 post, read on 2026-10-02. The Sume steps use video inspect, video trim and Timeline 1.0. Sume has no one-call silence remover; this is a recipe, and the cuts are yours to check.

What did YouTube say the editing assistant will do?

The post describes a new conversational editing tool across Shorts and the YouTube Create app that guides you through every edit: start from suggested prompts or your own for a first draft, then keep chatting to choose from suggestions to automatically trim pauses, reorder for better pacing and more. The availability is listed as coming soon, with no date.

For an event recap, that is the right idea: a talk with long stage silences and applause dead air gets tighter. If you need it before the tool ships, or for a long file Shorts would not take, build it from parts.

How do you find the gaps?

Call POST /v1/video-inspect with frames: false, transcribe: true and segmentation: { "mode": "sentence", "silence_split_seconds": 0.8 }. The docs say sentence mode returns gapless segments[], and silence_split_seconds accepts 0.2 to 3, so a lower number splits at shorter silences. Transcription is $0.01 per audio minute, the hint is capped at 600 seconds, and a clip with no audio fails inspect_source_has_no_audio.

Because the segments are gapless by design, you may not see a literal gap in them; the useful signal is where the speaker stopped, so also compare each segment's length with the words it holds. The script below shows the idea on a list of start and end times. It assumes each segment has start and end in seconds; check your own response and adjust the field names.

segments = [(0.0, 6.2), (6.3, 11.0), (17.8, 22.4), (22.5, 31.0)]
GAP = 2.0

keep, start = [], segments[0][0]
for (a0, a1), (b0, b1) in zip(segments, segments[1:]):
    if b0 - a1 >= GAP:
        keep.append((start, a1))
        start = b0
keep.append((start, segments[-1][1]))

for s, e in keep:
    print(f"trim start={s} duration={round(e - s, 2)}")

How do you cut and rejoin the keepers?

Send one POST /v1/video-trim per keeper range, each with its own Idempotency-Key; $0.02 each. Then render them in order with POST /v1/timeline-1.0/render. Timeline needs an audio spine: detach the original audio with audio detach and trim the matching ranges, or pass audio.parts[] slices so the sound follows the picture. Plan first with the unbilled POST /v1/timeline-1.0/plan.

Add a short transition of type fade of 0.25 seconds on each slot after the first to hide the jump, but note that more than eight adjacent fades is refused as too_many_chained_transitions; use a hard cut for some seams.

Treat the output as a rough cut. Cuts land on sentence boundaries, so most keeper ranges start and end cleanly, but a laugh or an audience reaction at the edge of a range may be lost. Widen each range by about 0.3 seconds at both ends when the speaker pauses for effect.

What can go wrong?

Cutting at a fixed silence length will clip a speaker's dramatic pause, a laugh or a slow joke. Review the result, and keep a gap of 0.5 to 1 second so the speech does not sound chopped. Speech-to-text may also mislabel applause as speech.

Source length matters: trim and inspect read at most 1800 seconds. A full-day conference track must be split first. For a looser recap built from phone clips instead of a single talk, see event recap from phone clips.

Check the joined result by listening to every seam. Two spoken sentences cut together can change the rhythm of a talk, and a fade hides a jump in the picture but not in the sound.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume