Find dead air in a Reel: sentence segments and silence split
Sume video inspect splits a transcript into sentence segments at pauses of 0.2 to 3 s. Use the gaps between segments to find dead air before you trim.

To find dead air in a Reel before you trim it, ask Sume's video inspect for a transcript with sentence segmentation, then look at the gaps between consecutive segments. The docs say segmentation.mode: "sentence" returns gapless, caption-line-shaped segments[], and that silence_split_seconds between 0.2 and 3 sets how long a pause has to be before the line breaks. A short value splits on small pauses; a long value splits only on the long ones. That makes the setting a dial for how much silence counts as dead air. Instagram's Reels page lists editing tools for trimming clips (read 2026-10-03), and this is a way to find what to trim, not a replacement for watching the cut.
Pick a threshold
Start with 1 second. A pause of a second or more in a talking-head Reel usually reads as a stumble, and shorter ones read as breath. Then widen or narrow it, depending on what you see. The cost does not change: the transcript is $0.01 per audio minute whatever the segmentation setting, and the docs say the segmentation fields without transcribe: true are refused with video_inspect_transcribe_required.
| Setting | Value | Effect |
|---|---|---|
| segmentation.mode | sentence | Adds gapless sentence segments |
| silence_split_seconds | 0.2 to 3 | Pause length that splits a line |
| transcribe | true | Required for either |
| frames | false | Skips stills so you pay for the transcript only |
Request the segments
The segment field names are not spelled out in the doc page I read, so the script prints the keys of the first segment and uses start and end if they are present. Run it once on a clip, read the output, and adjust the keys to match what comes back.
import json, os, urllib.request
API = "https://api.sume.com/v1"
def post(path, body, key):
req = urllib.request.Request(
f"{API}{path}",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": key,
},
method="POST",
)
with urllib.request.urlopen(req) as res:
return json.load(res)
res = post(
"/video-inspect",
{
"video_url": os.environ["SUME_CLIP_URL"],
"frames": False,
"transcribe": True,
"language_code": "en",
"segmentation": {"mode": "sentence", "silence_split_seconds": 1.0},
},
"dead-air-segments-v1",
)
inspect = res.get("video_inspect", res)
segments = (inspect.get("transcript") or {}).get("segments", [])
print(len(segments), "segments")
if segments:
print("keys:", sorted(segments[0]))
for before, after in zip(segments, segments[1:]):
if "end" in before and "start" in after:
gap = after["start"] - before["end"]
print(f"gap {gap:.2f}s after {before['end']:.2f}s")
From gaps to cuts
A gap list tells you where; a cut is still your decision. For each long gap, decide whether the pause is dead air or a deliberate beat before a punchline. Then cut the keep-ranges with video trim, which takes start plus exactly one of end or duration and returns a new MP4 for $0.02. Join the pieces in a Timeline render, which also accepts transitions of up to a second.
Trimming at sentence boundaries sounds cleaner than trimming mid-word, but it can clip breath sounds, so keep a little room on either side. The exact precision default is frame-accurate. If your clip has no audio, the inspect call fails with inspect_source_has_no_audio; check probe.has_audio first.
Finally, watch the result once. A script cannot tell you whether a pause was funny.
What this does not find
Silence in a transcript is a gap between words, so it misses dead air that has background noise, music or a long on-screen pause with narration over it. It also misses visual dead time, such as a locked-off shot where nothing moves while someone is talking. For that, pull stills with video frames at the times in question and look.
The segmentation also depends on the transcript, which depends on the audio. A noisy clip may split lines in odd places or merge two sentences, so a gap list is a starting point, not a verdict. Read it against the video the first few times, and only automate the cuts once you trust it on your footage.
A useful habit is to keep the threshold with the clip. Record the silence_split_seconds value you used next to the cut list, because the same footage run at 0.5 and at 2 gives different gap lists, and you will want to reproduce the one you trusted. Because the transcript is billed per audio minute, rerunning at a new threshold costs a cent or two for a short clip, which makes tuning cheap compared with a wrong cut.
Short Reels have the most to gain. In a 30-second clip, a single 1.5 second pause is five percent of the runtime, and removing two of them tightens the opening noticeably. In a ten-minute recording the same pauses are noise, so set the threshold higher for long material and lower for short.
Sources
Related posts
More in Developers
- Find every Sora call left in your codebase: a retired-id scan
OpenAI's Videos API shut down 2026-09-24. A 20-line Python scan lists every file still naming a retired sora-2 model id so you can swap them.
- Find micro-drama premises: TikTok trending search for $0.10 a call
One trending-videos search returns ranked TikTok metadata for a keyword, from 1 to 50 results, for $0.10. How to read it for premises, not to copy clips.
- Five Sume video submit errors and which ones to retry
insufficient_credits, idempotency_conflict, queue_full, rate_limited and provider_capacity_exceeded each need a different response. What to do, in a table.
- Flask webhook receiver for Sume: verify sume-v1, refuse empty secret
A Flask route that verifies the Sume signature on the raw body, takes either rotation entry, checks the replay window, and will not boot without a secret.
Written by Sume