Speaking rate in words per minute from Sume STT word times (Python)
Compute words per minute for a recording from the words[] start and end times Sume STT returns, plus a per-minute pacing table. Offline Python, no API call.

Short answer
Words per minute is the number of words divided by the time between the first word's start and the last word's end, times 60. Sume STT returns words[] as {word, start, end} with every transcript, so no extra call is needed and the script below runs offline on that array.
The transcript shape is documented in the video inspect page: transcript holds text, words[] and an optional sentence segments[].
Overall rate and rate without long pauses
Two figures are useful and they differ. The overall rate counts the pauses between phrases, so it is lower. The speaking rate inside phrases is higher, because it leaves out gaps longer than a threshold. A voiceover coach, a subtitle reader limit and a TTS speed setting each care about a different one.
The sample words below are invented to show the shape. Replace the list with transcript.words.
def wpm(words, gap=0.6):
if not words:
return 0.0, 0.0
span = words[-1]["end"] - words[0]["start"]
talk = sum(w["end"] - w["start"] for w in words)
for a, b in zip(words, words[1:]):
pause = b["start"] - a["end"]
if pause <= gap:
talk += pause
overall = len(words) / span * 60
inside = len(words) / talk * 60
return round(overall), round(inside)
words = [
{"word": "Welcome", "start": 0.0, "end": 0.5},
{"word": "to", "start": 0.5, "end": 0.6},
{"word": "the", "start": 0.6, "end": 0.7},
{"word": "demo", "start": 0.7, "end": 1.1},
{"word": "Today", "start": 2.3, "end": 2.7},
{"word": "we", "start": 2.7, "end": 2.8},
{"word": "build", "start": 2.8, "end": 3.2},
]
print(wpm(words))Reading the output
The gap argument is yours to set. At 0.6 seconds the pause between "demo" and "Today" (1.2 s) is treated as a break and left out of the inside-phrase figure.
Slice words by time to get a pacing table, one row per minute of audio. Use the second a word starts to choose its minute, so a word is counted once. A row that is far above or below the rest points at a rushed or dragging section.
| Minute | Words starting in the minute | Words per minute |
|---|---|---|
| 0-1 | count of words with start in [0, 60) | that count |
| 1-2 | count of words with start in [60, 120) | that count |
| 2-3 | count of words with start in [120, 180) | that count |
What the number can and cannot tell you
Word times are model output and are not frame-exact, so a result within a few words per minute is noise. Compare recordings made the same way.
Cost and limits
STT is $0.01 per audio minute at the public rate, so a ten-minute recording costs about 10 cents to measure. Maximum audio per request is 10 minutes; for a longer recording, split it first and add the offset of each part to its word times before you run the script.
Using the rate with TTS
If you plan to match a recorded pace with synthetic speech, TTS has a generation_config.speed multiplier from 0.6 to 1.5, so a measured rate can guide the first value to try.
Choosing a target rate
The code gives you your own rate; the target comes from the job. A voiceover for a short ad usually runs faster than a meditation, and a training clip for new learners slower than a news read. Do not borrow a number from a chart; record one recording you consider good, run the script on it, and use that rate as your reference.
Compare the inside-phrase figure across recordings of the same kind. If one speaker is much faster inside phrases but similar overall, they are rushing words and then pausing, which is a different fix from simply speaking slower.
Words are counted as they appear in the transcript, so numbers, abbreviations and hyphenated terms can split differently from how you would count. Use the same transcription path for every recording you compare.
- Record a reference you like, and measure it first.
- Compare like with like: same language, same type of content.
- Keep one counting method.
Sources
Related posts
More in Developers
- Speech to text API in Go: transcribe audio with net/http
Transcribe audio in Go using only the standard library: submit to Sume STT, poll the job and print the text. A 30-line program at one cent per audio minute.
- Speech to text API in Node.js: transcribe audio with fetch
Transcribe audio in Node.js with built-in fetch: submit to Sume STT, poll the job and print sentence segments with timestamps. 30 lines, no dependencies.
- Speech to text API in Ruby: transcribe audio with Net::HTTP
Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.
- Speech to text API in Swift: transcribe audio with URLSession
Transcribe audio in Swift with async URLSession: submit to Sume STT, poll the job, print text. No packages, 28 lines, one cent per audio minute of audio.
Written by Sume