Turn STT words into paragraphs: break at pauses over 1.2 seconds
Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.

A transcript as one block of text is hard to read. The sentence segments[] from Sume STT help, but a talk has a bigger structure than sentences: a speaker finishes a point, breathes, and starts the next. That shows up in the word times as a gap. Because words[] carries start and end for every token, you can find the gaps yourself and cut paragraphs there. It needs no extra API call, so it costs nothing beyond the transcription.
The rule
For each pair of neighboring words, the gap is the next word's start minus the previous word's end. Start a new paragraph when the gap exceeds a threshold, and the paragraph already has at least a few sentences. A threshold of 1.2 seconds is a starting point for a calm talk. For quick conversation, 0.8 seconds. Tune it on one file you know well.
def paragraphs(words, gap=1.2, min_words=25):
paras, cur, prev_end = [], [], None
for w in words:
if w.get("type") != "word":
continue
long_pause = prev_end is not None and w["start"] - prev_end > gap
if long_pause and len(cur) >= min_words:
paras.append(" ".join(cur))
cur = []
cur.append(w["word"])
prev_end = w["end"]
if cur:
paras.append(" ".join(cur))
return paras
# words = result["words"] from a completed /v1/stt-1.0/transcribe job
words = [{"word": "Hello", "type": "word", "start": 0.0, "end": 0.4}]
print("\n\n".join(paragraphs(words)))Why use the words and not the text
The text field is a string with no timing. The words list also holds spacing tokens, which the code skips, so your paragraphs hold only spoken words. Joining with single spaces puts spaces before punctuation if the provider returns punctuation as its own token; check one result before you trust the join, and trim , to , if you see it.
Tuning
- Raise
gapwhen paragraphs come out too short. - Raise
min_wordsto avoid one-sentence paragraphs after a pause for effect. - Print the gap sizes of one file sorted descending. The top ten are usually the real topic changes.
- Add the first word's time as a prefix for each paragraph, such as
[04:12], to get a navigable transcript.
Cost
The transcription is about $0.01 per audio minute, 10 cents for a full 10-minute job. The paragraph step is local code. For files over 10 minutes, join the word lists after you add each part's offset, then run this over the combined list.
Sources
Related posts
More in Developers
- AI music API in TypeScript: generate, wait and save an MP3
Call Sume's Music Router from TypeScript: generateMusicRouter, waitForJob, then download the audio artifact. $0.125 per track, Google lists Lyria 3.5 at $0.08.
- TypeScript exhaustive switch over a Sume run's terminal status
A run ends as completed, failed, canceled or skipped, and the last two send no webhook. Use a never check so a new status fails the build.
- TypeScript webhook verifier for Wan 3.0 clips: refuse an empty secret
A TypeScript verifier for Sume webhooks: HMAC SHA 256 over timestamp.body, rotation entries, a 5 minute replay window and a hard refusal of an empty secret.
- Ukrainian text to speech API: set language uk and price a script
Cartesia Sonic 3.6 lists Ukrainian. Send language uk to Sume TTS 1.0, handle the 409 voice mismatch, and price a 5,000-character script at $0.24.
Written by Sume