Trim silence from a voiceover using STT word timings (Python)

Find the dead air in a voiceover from Sume STT words[] timings: list gaps over a threshold in Python, then cut with timeline audio ranges.

4 min readSume
All posts

Silence in a voiceover shows up as gaps between words. Sume STT 1.0 always returns words[] with word, start and end in seconds from the start of the audio, so you can find every pause longer than a threshold without any audio library. Then you cut around them with ranges.

Find the gaps

The script below runs on a sample words list shaped like the real result. Swap the literal for result["words"] from your STT job. It prints each gap above 0.6 seconds and builds the keep ranges.

words = [
    {"word": "Hello", "start": 0.10, "end": 0.45},
    {"word": "and", "start": 0.50, "end": 0.62},
    {"word": "welcome", "start": 1.90, "end": 2.40},
    {"word": "back", "start": 2.45, "end": 2.80},
]
MAX_GAP = 0.6
keep, begin = [], words[0]["start"]
for prev, nxt in zip(words, words[1:]):
    gap = nxt["start"] - prev["end"]
    if gap > MAX_GAP:
        print(f"gap {gap:.2f}s after {prev['word']!r}")
        keep.append({"start": begin, "end": prev["end"] + 0.07})
        begin = nxt["start"] - 0.07
keep.append({"start": begin, "end": words[-1]["end"]})
print(keep)

Why a margin

The 0.07 second margin mirrors the 70 ms boundary lead that Sume uses for sentence cuts, so a cut does not clip the last consonant. The docs call the default boundary_lead_ms 70, and it can be 0 to 500.

Cut and join

Cut the file with timeline audio: operation: "split" takes a url and ranges[] of up to 20, each { start, end }. Then operation: "concat" joins up to 20 parts into one gapless file with no re-synthesis. If you have more than 20 keep ranges, run the split in batches. The input URLs must be your workspace's media.sume.com audio.

Limits for this recipe, read 2026-10-06:

Silence trim, limits, read 2026-10-06
StepSurfaceLimit
TranscribePOST /v1/stt-1.0/transcribe600 s per file
SplitPOST /v1/timeline-1.0/audio1 to 20 ranges
JoinPOST /v1/timeline-1.0/audio1 to 20 parts

Where this recipe fails

A word timing is the model's estimate. On a fast read the gap between two words can be a few hundredths of a second, so a threshold below about 0.3 seconds starts to cut breaths and natural rhythm. Keep the limit high for narration and lower it only for a deliberately clipped ad voice.

Also check the start and the end. Leading and trailing silence is not a gap between two words, so the script above keeps the first word's start and the last word's end. Add a short pad if a platform adds its own fade.

Tighten the threshold for fast ads and loosen it for narration, since a breath is not dead air. Always listen to the joined file, because word timings come from a model and can be off by a few hundredths of a second.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume