Find the timestamp of a phrase in a recording with Sume STT words[]

Every Sume STT job returns words[] with start and end seconds. A short Python function finds where a phrase is said, so you can jump to or cut the exact moment.

5 min readSume
All posts

The answer

Run the recording through POST /v1/stt-1.0/transcribe, then search the returned words[] for the phrase. Each item is {word, start, end} in seconds from the start of the audio, so the phrase starts at the first word's start and ends at the last word's end. There is no flag to enable word times; they always come back.

That makes a spoken search tool a ten-line function. The cost is $0.01 per audio minute, so a 10-minute recording, the longest a single job takes, costs $0.10.

What comes back

The shape of the result matters more than the search.

Sume STT 1.0 request and result fields used here (read 2026-10-04)
FieldWhereMeaning
audio_urlRequestPublic HTTPS audio URL; prefer a Sume media URL
duration_secondsRequest1 to 600; omit and one minute is reserved
language_codeRequestOptional hint; omit for auto-detect
words[]ResultEach word with start and end in seconds
textResultThe full transcript

The function

Normalise punctuation and case before you compare, because the transcript carries punctuation and your search phrase will not. The function returns the first match, or None.

import re

def norm(w):
    return re.sub(r"[^\w']", "", w.lower())

def find_phrase(words, phrase):
    want = [norm(t) for t in phrase.split()]
    have = [norm(w["word"]) for w in words]
    for i in range(len(have) - len(want) + 1):
        if have[i:i + len(want)] == want:
            return words[i]["start"], words[i + len(want) - 1]["end"]
    return None

sample = [
    {"word": "Thanks,", "start": 1.2, "end": 1.5},
    {"word": "everyone.", "start": 1.5, "end": 2.0},
    {"word": "Our", "start": 2.4, "end": 2.6},
    {"word": "Q4", "start": 2.6, "end": 3.0},
]
print(find_phrase(sample, "our q4"))   # (2.4, 3.0)

Using the time

Once you have the start and end, you can seek a player to it, or cut the clip. For a video, trim the range with Sume's video trim, or detach just that range of audio with range: {start, end}. The audio detach docs describe the range field. Add a half second of padding on both sides so the cut does not clip a breath or the end of a word.

If the file is longer than 10 minutes, split it, transcribe each piece and add the piece's offset to every time, as in the 30-minute plan. A phrase that crosses a boundary will be missed, so overlap the pieces by a few seconds.

  • Match on words, not characters, to avoid cutting mid-word.
  • Search for a few words rather than one; common words match everywhere.
  • Check the match by listening before you cut anything public.
  • Keep the job id with the timestamp so you can re-fetch the words.

Worked example: finding a quote for a clip

A marketer wants the moment a speaker says "our Q4 numbers" in a 9-minute webinar. She transcribes the recording with duration_seconds: 540, which costs 9 x $0.01 = $0.09. The function returns, say, 312.4 to 314.1 seconds. She pads one half second each side, so the range is 311.9 to 314.6, and detaches or trims that range.

Run the search on several phrases from one transcript. The STT job is paid once; every search afterwards is free, since you are only reading words[]. Store the words array in your own database next to the job id, and a whole library of recordings becomes searchable by spoken phrase without touching the audio again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume