Find dead air in a Reel: sentence segments and silence split

Sume video inspect splits a transcript into sentence segments at pauses of 0.2 to 3 s. Use the gaps between segments to find dead air before you trim.

4 min readSume
All posts

To find dead air in a Reel before you trim it, ask Sume's video inspect for a transcript with sentence segmentation, then look at the gaps between consecutive segments. The docs say segmentation.mode: "sentence" returns gapless, caption-line-shaped segments[], and that silence_split_seconds between 0.2 and 3 sets how long a pause has to be before the line breaks. A short value splits on small pauses; a long value splits only on the long ones. That makes the setting a dial for how much silence counts as dead air. Instagram's Reels page lists editing tools for trimming clips (read 2026-10-03), and this is a way to find what to trim, not a replacement for watching the cut.

Pick a threshold

Start with 1 second. A pause of a second or more in a talking-head Reel usually reads as a stumble, and shorter ones read as breath. Then widen or narrow it, depending on what you see. The cost does not change: the transcript is $0.01 per audio minute whatever the segmentation setting, and the docs say the segmentation fields without transcribe: true are refused with video_inspect_transcribe_required.

Segmentation settings, read 2026-10-03
SettingValueEffect
segmentation.modesentenceAdds gapless sentence segments
silence_split_seconds0.2 to 3Pause length that splits a line
transcribetrueRequired for either
framesfalseSkips stills so you pay for the transcript only

Request the segments

The segment field names are not spelled out in the doc page I read, so the script prints the keys of the first segment and uses start and end if they are present. Run it once on a clip, read the output, and adjust the keys to match what comes back.

import json, os, urllib.request

API = "https://api.sume.com/v1"

def post(path, body, key):
    req = urllib.request.Request(
        f"{API}{path}",
        data=json.dumps(body).encode(),
        headers={
            "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
            "Content-Type": "application/json",
            "Idempotency-Key": key,
        },
        method="POST",
    )
    with urllib.request.urlopen(req) as res:
        return json.load(res)

res = post(
    "/video-inspect",
    {
        "video_url": os.environ["SUME_CLIP_URL"],
        "frames": False,
        "transcribe": True,
        "language_code": "en",
        "segmentation": {"mode": "sentence", "silence_split_seconds": 1.0},
    },
    "dead-air-segments-v1",
)
inspect = res.get("video_inspect", res)
segments = (inspect.get("transcript") or {}).get("segments", [])
print(len(segments), "segments")
if segments:
    print("keys:", sorted(segments[0]))
for before, after in zip(segments, segments[1:]):
    if "end" in before and "start" in after:
        gap = after["start"] - before["end"]
        print(f"gap {gap:.2f}s after {before['end']:.2f}s")

From gaps to cuts

A gap list tells you where; a cut is still your decision. For each long gap, decide whether the pause is dead air or a deliberate beat before a punchline. Then cut the keep-ranges with video trim, which takes start plus exactly one of end or duration and returns a new MP4 for $0.02. Join the pieces in a Timeline render, which also accepts transitions of up to a second.

Trimming at sentence boundaries sounds cleaner than trimming mid-word, but it can clip breath sounds, so keep a little room on either side. The exact precision default is frame-accurate. If your clip has no audio, the inspect call fails with inspect_source_has_no_audio; check probe.has_audio first.

Finally, watch the result once. A script cannot tell you whether a pause was funny.

What this does not find

Silence in a transcript is a gap between words, so it misses dead air that has background noise, music or a long on-screen pause with narration over it. It also misses visual dead time, such as a locked-off shot where nothing moves while someone is talking. For that, pull stills with video frames at the times in question and look.

The segmentation also depends on the transcript, which depends on the audio. A noisy clip may split lines in odd places or merge two sentences, so a gap list is a starting point, not a verdict. Read it against the video the first few times, and only automate the cuts once you trust it on your footage.

A useful habit is to keep the threshold with the clip. Record the silence_split_seconds value you used next to the cut list, because the same footage run at 0.5 and at 2 gives different gap lists, and you will want to reproduce the one you trusted. Because the transcript is billed per audio minute, rerunning at a new threshold costs a cent or two for a short clip, which makes tuning cheap compared with a wrong cut.

Short Reels have the most to gain. In a 30-second clip, a single 1.5 second pause is five percent of the runtime, and removing two of them tightens the opening noticeably. In a ten-minute recording the same pauses are noise, so set the threshold higher for long material and lower for short.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume