Find dead air in a voice-agent call recording with STT word gaps

Long silences hurt voice-agent calls. Transcribe a recording with Sume STT and list every pause over a threshold, with times. Python sketch.

4 min readSume
All posts

Transcribe the recording with Sume STT, keep only the entries of type word, and subtract each word's end from the next word's start. Any gap above your threshold is dead air, and the result gives you the second it began.

This finds pauses, not their cause. The documented STT result has word, start, end and type per entry, and no speaker field, so you cannot tell from the data whether the agent or the caller went quiet.

Why measure silence in agent calls?

Voice agents are where a pause is most noticeable. Google's Gemini API changelog lists Gemini 3.8 Live as generally available on 2026-09-15, and dailyaipedia reports ElevenLabs added call queueing with hold audio on 2026-09-14. Whatever stack you use, the recording is the evidence: it shows how long the caller waited, not how long the vendor says a turn takes.

Sume's STT is a batch job on a recording, not a live stream, so this is an after-the-call audit.

How do I get the words?

POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url and always returns words[] with start and end in seconds. If your recording is a video on media.sume.com, use audio detach first with mono and 16 kHz, which the docs call the STT shape. Pass duration_seconds (up to 600) and split longer calls into parts, adding each part's offset back to its times.

Entries have a type. Skip the ones with type spacing, or a space token will hide a real pause.

Fields used from the STT result (Sume OpenAPI, read 2026-10-02).
FieldMeaning
words[].wordToken text
words[].startSeconds from the audio start
words[].endSeconds from the audio start
words[].typeword or spacing, when supplied

What does the gap finder look like?

The sample list below mimics a result with one spacing token and two long pauses. Swap it for the real words array. The 1.5 second threshold is an example; pick what suits your calls.

words = [  # shape of result.words from an STT job: word, start, end, type
    {"word": "Thanks", "start": 0.0, "end": 0.4, "type": "word"},
    {"word": " ", "start": 0.4, "end": 0.5, "type": "spacing"},
    {"word": "for", "start": 0.5, "end": 0.7, "type": "word"},
    {"word": "calling", "start": 0.7, "end": 1.2, "type": "word"},
    {"word": "Let", "start": 4.1, "end": 4.3, "type": "word"},
    {"word": "me", "start": 4.3, "end": 4.4, "type": "word"},
    {"word": "check", "start": 9.0, "end": 9.5, "type": "word"},
]
THRESHOLD = 1.5  # seconds; choose your own
spoken = [w for w in words if w.get("type", "word") == "word"]
for prev, nxt in zip(spoken, spoken[1:]):
    gap = nxt["start"] - prev["end"]
    if gap >= THRESHOLD:
        print(f"{gap:.1f}s gap after '{prev['word']}' at {prev['end']:.1f}s")

How do I read the output?

Open the recording at each reported time. Hold music, a transfer or a caller thinking all show as a gap, and so does real dead air. Count gaps per call and compare the biggest ones across a day of calls. If you also need to know who was speaking, label the turns yourself from the audio, since the STT result does not do it.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume