Speaking rate in words per minute from Sume STT word times (Python)

Compute words per minute for a recording from the words[] start and end times Sume STT returns, plus a per-minute pacing table. Offline Python, no API call.

5 min readSume
All posts

Short answer

Words per minute is the number of words divided by the time between the first word's start and the last word's end, times 60. Sume STT returns words[] as {word, start, end} with every transcript, so no extra call is needed and the script below runs offline on that array.

The transcript shape is documented in the video inspect page: transcript holds text, words[] and an optional sentence segments[].

Overall rate and rate without long pauses

Two figures are useful and they differ. The overall rate counts the pauses between phrases, so it is lower. The speaking rate inside phrases is higher, because it leaves out gaps longer than a threshold. A voiceover coach, a subtitle reader limit and a TTS speed setting each care about a different one.

The sample words below are invented to show the shape. Replace the list with transcript.words.

def wpm(words, gap=0.6):
    if not words:
        return 0.0, 0.0
    span = words[-1]["end"] - words[0]["start"]
    talk = sum(w["end"] - w["start"] for w in words)
    for a, b in zip(words, words[1:]):
        pause = b["start"] - a["end"]
        if pause <= gap:
            talk += pause
    overall = len(words) / span * 60
    inside = len(words) / talk * 60
    return round(overall), round(inside)

words = [
    {"word": "Welcome", "start": 0.0, "end": 0.5},
    {"word": "to", "start": 0.5, "end": 0.6},
    {"word": "the", "start": 0.6, "end": 0.7},
    {"word": "demo", "start": 0.7, "end": 1.1},
    {"word": "Today", "start": 2.3, "end": 2.7},
    {"word": "we", "start": 2.7, "end": 2.8},
    {"word": "build", "start": 2.8, "end": 3.2},
]
print(wpm(words))

Reading the output

The gap argument is yours to set. At 0.6 seconds the pause between "demo" and "Today" (1.2 s) is treated as a break and left out of the inside-phrase figure.

Slice words by time to get a pacing table, one row per minute of audio. Use the second a word starts to choose its minute, so a word is counted once. A row that is far above or below the rest points at a rushed or dragging section.

Pacing table layout (illustrative layout, no measured data)
MinuteWords starting in the minuteWords per minute
0-1count of words with start in [0, 60)that count
1-2count of words with start in [60, 120)that count
2-3count of words with start in [120, 180)that count

What the number can and cannot tell you

Word times are model output and are not frame-exact, so a result within a few words per minute is noise. Compare recordings made the same way.

Cost and limits

STT is $0.01 per audio minute at the public rate, so a ten-minute recording costs about 10 cents to measure. Maximum audio per request is 10 minutes; for a longer recording, split it first and add the offset of each part to its word times before you run the script.

Using the rate with TTS

If you plan to match a recorded pace with synthetic speech, TTS has a generation_config.speed multiplier from 0.6 to 1.5, so a measured rate can guide the first value to try.

Choosing a target rate

The code gives you your own rate; the target comes from the job. A voiceover for a short ad usually runs faster than a meditation, and a training clip for new learners slower than a news read. Do not borrow a number from a chart; record one recording you consider good, run the script on it, and use that rate as your reference.

Compare the inside-phrase figure across recordings of the same kind. If one speaker is much faster inside phrases but similar overall, they are rushing words and then pausing, which is a different fix from simply speaking slower.

Words are counted as they appear in the transcript, so numbers, abbreviations and hyphenated terms can split differently from how you would count. Use the same transcription path for every recording you compare.

  • Record a reference you like, and measure it first.
  • Compare like with like: same language, same type of content.
  • Keep one counting method.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume