Find dead air in a voice-agent call recording with STT word gaps
Long silences hurt voice-agent calls. Transcribe a recording with Sume STT and list every pause over a threshold, with times. Python sketch.

Transcribe the recording with Sume STT, keep only the entries of type word, and subtract each word's end from the next word's start. Any gap above your threshold is dead air, and the result gives you the second it began.
This finds pauses, not their cause. The documented STT result has word, start, end and type per entry, and no speaker field, so you cannot tell from the data whether the agent or the caller went quiet.
Why measure silence in agent calls?
Voice agents are where a pause is most noticeable. Google's Gemini API changelog lists Gemini 3.8 Live as generally available on 2026-09-15, and dailyaipedia reports ElevenLabs added call queueing with hold audio on 2026-09-14. Whatever stack you use, the recording is the evidence: it shows how long the caller waited, not how long the vendor says a turn takes.
Sume's STT is a batch job on a recording, not a live stream, so this is an after-the-call audit.
How do I get the words?
POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url and always returns words[] with start and end in seconds. If your recording is a video on media.sume.com, use audio detach first with mono and 16 kHz, which the docs call the STT shape. Pass duration_seconds (up to 600) and split longer calls into parts, adding each part's offset back to its times.
Entries have a type. Skip the ones with type spacing, or a space token will hide a real pause.
| Field | Meaning |
|---|---|
| words[].word | Token text |
| words[].start | Seconds from the audio start |
| words[].end | Seconds from the audio start |
| words[].type | word or spacing, when supplied |
What does the gap finder look like?
The sample list below mimics a result with one spacing token and two long pauses. Swap it for the real words array. The 1.5 second threshold is an example; pick what suits your calls.
words = [ # shape of result.words from an STT job: word, start, end, type
{"word": "Thanks", "start": 0.0, "end": 0.4, "type": "word"},
{"word": " ", "start": 0.4, "end": 0.5, "type": "spacing"},
{"word": "for", "start": 0.5, "end": 0.7, "type": "word"},
{"word": "calling", "start": 0.7, "end": 1.2, "type": "word"},
{"word": "Let", "start": 4.1, "end": 4.3, "type": "word"},
{"word": "me", "start": 4.3, "end": 4.4, "type": "word"},
{"word": "check", "start": 9.0, "end": 9.5, "type": "word"},
]
THRESHOLD = 1.5 # seconds; choose your own
spoken = [w for w in words if w.get("type", "word") == "word"]
for prev, nxt in zip(spoken, spoken[1:]):
gap = nxt["start"] - prev["end"]
if gap >= THRESHOLD:
print(f"{gap:.1f}s gap after '{prev['word']}' at {prev['end']:.1f}s")
How do I read the output?
Open the recording at each reported time. Hold music, a transfer or a caller thinking all show as a gap, and so does real dead air. Count gaps per call and compare the biggest ones across a day of calls. If you also need to know who was speaking, label the turns yourself from the audio, since the STT result does not do it.
Sources
Related posts
More in Use cases
- Flow's 50 daily credits don't roll over: plan a production week
Flow refreshes 50 free credits a day after first use and drops what you skip. How to plan a week of clips, and how a Sume USD balance behaves instead.
- Flow monthly credits by Google AI plan: how many clips fit
Flow gives 200, 1,000, 10,000 or 25,000 bonus credits a month by plan. What that buys in Veo 3.1 and Gemini Omni Flash clips, and how a Sume balance differs.
- FTC Disclosures 101: why 'sp' fails and what to burn into a video
The FTC calls 'sp' and 'spon' unclear and wants video disclosures in audio and visuals. Plain wording to burn into an avatar clip with Sume caption cues.
- FTC fake reviews rule: founder avatar testimonial and insider reviews
The FTC's fake reviews rule covers AI reviewers who do not exist and insider reviews with no disclosed connection. What that means for an avatar testimonial.
Written by Sume