Turn STT words into paragraphs: break at pauses over 1.2 seconds

Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.

4 min readSume
All posts

A transcript as one block of text is hard to read. The sentence segments[] from Sume STT help, but a talk has a bigger structure than sentences: a speaker finishes a point, breathes, and starts the next. That shows up in the word times as a gap. Because words[] carries start and end for every token, you can find the gaps yourself and cut paragraphs there. It needs no extra API call, so it costs nothing beyond the transcription.

The rule

For each pair of neighboring words, the gap is the next word's start minus the previous word's end. Start a new paragraph when the gap exceeds a threshold, and the paragraph already has at least a few sentences. A threshold of 1.2 seconds is a starting point for a calm talk. For quick conversation, 0.8 seconds. Tune it on one file you know well.

def paragraphs(words, gap=1.2, min_words=25):
    paras, cur, prev_end = [], [], None
    for w in words:
        if w.get("type") != "word":
            continue
        long_pause = prev_end is not None and w["start"] - prev_end > gap
        if long_pause and len(cur) >= min_words:
            paras.append(" ".join(cur))
            cur = []
        cur.append(w["word"])
        prev_end = w["end"]
    if cur:
        paras.append(" ".join(cur))
    return paras

# words = result["words"] from a completed /v1/stt-1.0/transcribe job
words = [{"word": "Hello", "type": "word", "start": 0.0, "end": 0.4}]
print("\n\n".join(paragraphs(words)))

Why use the words and not the text

The text field is a string with no timing. The words list also holds spacing tokens, which the code skips, so your paragraphs hold only spoken words. Joining with single spaces puts spaces before punctuation if the provider returns punctuation as its own token; check one result before you trust the join, and trim , to , if you see it.

Tuning

  • Raise gap when paragraphs come out too short.
  • Raise min_words to avoid one-sentence paragraphs after a pause for effect.
  • Print the gap sizes of one file sorted descending. The top ten are usually the real topic changes.
  • Add the first word's time as a prefix for each paragraph, such as [04:12], to get a navigable transcript.

Cost

The transcription is about $0.01 per audio minute, 10 cents for a full 10-minute job. The paragraph step is local code. For files over 10 minutes, join the word lists after you add each part's offset, then run this over the combined list.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume