Trim silence from a voiceover using STT word timings (Python)
Find the dead air in a voiceover from Sume STT words[] timings: list gaps over a threshold in Python, then cut with timeline audio ranges.

Silence in a voiceover shows up as gaps between words. Sume STT 1.0 always returns words[] with word, start and end in seconds from the start of the audio, so you can find every pause longer than a threshold without any audio library. Then you cut around them with ranges.
Find the gaps
The script below runs on a sample words list shaped like the real result. Swap the literal for result["words"] from your STT job. It prints each gap above 0.6 seconds and builds the keep ranges.
words = [
{"word": "Hello", "start": 0.10, "end": 0.45},
{"word": "and", "start": 0.50, "end": 0.62},
{"word": "welcome", "start": 1.90, "end": 2.40},
{"word": "back", "start": 2.45, "end": 2.80},
]
MAX_GAP = 0.6
keep, begin = [], words[0]["start"]
for prev, nxt in zip(words, words[1:]):
gap = nxt["start"] - prev["end"]
if gap > MAX_GAP:
print(f"gap {gap:.2f}s after {prev['word']!r}")
keep.append({"start": begin, "end": prev["end"] + 0.07})
begin = nxt["start"] - 0.07
keep.append({"start": begin, "end": words[-1]["end"]})
print(keep)Why a margin
The 0.07 second margin mirrors the 70 ms boundary lead that Sume uses for sentence cuts, so a cut does not clip the last consonant. The docs call the default boundary_lead_ms 70, and it can be 0 to 500.
Cut and join
Cut the file with timeline audio: operation: "split" takes a url and ranges[] of up to 20, each { start, end }. Then operation: "concat" joins up to 20 parts into one gapless file with no re-synthesis. If you have more than 20 keep ranges, run the split in batches. The input URLs must be your workspace's media.sume.com audio.
Limits for this recipe, read 2026-10-06:
| Step | Surface | Limit |
|---|---|---|
| Transcribe | POST /v1/stt-1.0/transcribe | 600 s per file |
| Split | POST /v1/timeline-1.0/audio | 1 to 20 ranges |
| Join | POST /v1/timeline-1.0/audio | 1 to 20 parts |
Where this recipe fails
A word timing is the model's estimate. On a fast read the gap between two words can be a few hundredths of a second, so a threshold below about 0.3 seconds starts to cut breaths and natural rhythm. Keep the limit high for narration and lower it only for a deliberately clipped ad voice.
Also check the start and the end. Leading and trailing silence is not a gap between two words, so the script above keeps the first word's start and the last word's end. Add a short pad if a platform adds its own fade.
Tighten the threshold for fast ads and loosen it for narration, since a breath is not dead air. Always listen to the joined file, because word timings come from a model and can be off by a few hundredths of a second.
Sources
Related posts
More in Developers
- Text-to-speech API in Node: submit, poll and save the MP3
A Node 18+ fetch example for the Sume TTS Router: submit with an Idempotency-Key, poll status_url, read the audio artifact and save an MP3 to disk.
- TTS model list API: read the catalog before you hardcode a Sonic id
Sume's TTS Router lists its models at GET /v1/tts-router/models. Read it, pin sonic-3.6, and treat sonic-preview as a beta channel that can change.
- tts_sentence_selection_invalid 422 on Sume TTS: what triggers it
Sume TTS returns 422 tts_sentence_selection_invalid for gaps, repeated jobs, unfinished jobs and partial coverage. Each cause and its fix.
- tts_source_integrity_mismatch 422: job differs from accepted script
verify-spine returns 422 tts_source_integrity_mismatch when a finished TTS job's text no longer matches the accepted script. What it checks and how to recover.
Written by Sume