Find the timestamp of a phrase in a recording with Sume STT words[]
Every Sume STT job returns words[] with start and end seconds. A short Python function finds where a phrase is said, so you can jump to or cut the exact moment.

The answer
Run the recording through POST /v1/stt-1.0/transcribe, then search the returned words[] for the phrase. Each item is {word, start, end} in seconds from the start of the audio, so the phrase starts at the first word's start and ends at the last word's end. There is no flag to enable word times; they always come back.
That makes a spoken search tool a ten-line function. The cost is $0.01 per audio minute, so a 10-minute recording, the longest a single job takes, costs $0.10.
What comes back
The shape of the result matters more than the search.
| Field | Where | Meaning |
|---|---|---|
| audio_url | Request | Public HTTPS audio URL; prefer a Sume media URL |
| duration_seconds | Request | 1 to 600; omit and one minute is reserved |
| language_code | Request | Optional hint; omit for auto-detect |
| words[] | Result | Each word with start and end in seconds |
| text | Result | The full transcript |
The function
Normalise punctuation and case before you compare, because the transcript carries punctuation and your search phrase will not. The function returns the first match, or None.
import re
def norm(w):
return re.sub(r"[^\w']", "", w.lower())
def find_phrase(words, phrase):
want = [norm(t) for t in phrase.split()]
have = [norm(w["word"]) for w in words]
for i in range(len(have) - len(want) + 1):
if have[i:i + len(want)] == want:
return words[i]["start"], words[i + len(want) - 1]["end"]
return None
sample = [
{"word": "Thanks,", "start": 1.2, "end": 1.5},
{"word": "everyone.", "start": 1.5, "end": 2.0},
{"word": "Our", "start": 2.4, "end": 2.6},
{"word": "Q4", "start": 2.6, "end": 3.0},
]
print(find_phrase(sample, "our q4")) # (2.4, 3.0)Using the time
Once you have the start and end, you can seek a player to it, or cut the clip. For a video, trim the range with Sume's video trim, or detach just that range of audio with range: {start, end}. The audio detach docs describe the range field. Add a half second of padding on both sides so the cut does not clip a breath or the end of a word.
If the file is longer than 10 minutes, split it, transcribe each piece and add the piece's offset to every time, as in the 30-minute plan. A phrase that crosses a boundary will be missed, so overlap the pieces by a few seconds.
- Match on words, not characters, to avoid cutting mid-word.
- Search for a few words rather than one; common words match everywhere.
- Check the match by listening before you cut anything public.
- Keep the job id with the timestamp so you can re-fetch the words.
Worked example: finding a quote for a clip
A marketer wants the moment a speaker says "our Q4 numbers" in a 9-minute webinar. She transcribes the recording with duration_seconds: 540, which costs 9 x $0.01 = $0.09. The function returns, say, 312.4 to 314.1 seconds. She pads one half second each side, so the range is 311.9 to 314.6, and detaches or trims that range.
Run the search on several phrases from one transcript. The STT job is paid once; every search afterwards is free, since you are only reading words[]. Store the words array in your own database next to the job id, and a whole library of recordings becomes searchable by spoken phrase without touching the audio again.
Sources
Related posts
More in Developers
- Find Veo 3.1 preview ids in your repo before Oct 22: a scanner
Google's Veo 3.1 preview ids shut down Oct 22, 2026. A short Python scanner lists every file and line using them, and a Sume id to switch to.
- Firestore create() with job_id as document id for Sume webhooks
Name the Firestore document after the Sume job_id and call create(). ALREADY_EXISTS marks a retry. Node Admin SDK sample, plus a claim state for crashes.
- Fit a voiceover to a 30-second slot with Sume TTS speed
Measure the first take, divide by the slot length, and set generation_config.speed (0.6 to 1.5). Why a big speed-up is better solved by cutting the script.
- Fit narration to a fixed slot: measure first, then set TTS speed
A 45-second cap or a 30-second slot decides your script. Render once with word timings, compute the speed ratio, and rewrite only if outside 0.6 to 1.5.
Written by Sume