Find the timestamp of a quote in a recording with Sume STT words
Sume's STT result returns every word with start and end seconds. A short Python function finds a quoted phrase and returns where to cut. About a cent a minute.

To find where a quote is spoken, transcribe the file with Sume STT 1.0 and search the returned words array, each entry {word, start, end} in seconds from the audio start. A match of the quote's words gives you the start of the first word and the end of the last. A 10-minute file costs 10 cents.
The function below is plain Python with no dependencies, and it was run against a small sample before publishing.
What the result contains
The result of a completed STT job has text, language_code, language_probability and words. Words are ordered by start, and the array is always present, even when empty. It is capped at 20,000 entries, which a 600-second transcript stays well under. Some tokens may be classed spacing in the optional type field, so skip those when matching.
STT 1.0 takes audio up to 600 seconds. For a longer recording, split it and add the slice offset to each match.
| Field | Meaning | Unit |
|---|---|---|
| words[].word | Token text | string |
| words[].start | Token start from audio start | seconds |
| words[].end | Token end | seconds |
| words[].type | Provider class, for example spacing | optional |
| words_truncated | Present only if capped at 20,000 | boolean |
The search function
Normalize both sides by lowercasing and stripping punctuation, then slide a window across the word list. Pad the cut by a fraction of a second so the clip does not clip the first syllable.
import re
def norm(s):
return re.sub(r"[^\w']+", "", s.lower())
def find_quote(words, quote, pad=0.2):
toks = [w for w in words if w.get("type") != "spacing"]
want = [norm(t) for t in quote.split() if norm(t)]
have = [norm(w["word"]) for w in toks]
for i in range(len(have) - len(want) + 1):
if have[i:i + len(want)] == want:
return (max(0.0, toks[i]["start"] - pad),
toks[i + len(want) - 1]["end"] + pad)
return None
sample = [{"word": "We", "start": 1.0, "end": 1.2},
{"word": "ship", "start": 1.2, "end": 1.6},
{"word": "on", "start": 1.6, "end": 1.7},
{"word": "Friday.", "start": 1.7, "end": 2.3}]
print(find_quote(sample, "ship on friday"))Then cut the clip
If the recording is a video you uploaded to Sume's media host, the trim route cuts a [start, end) range in seconds, and the quote's start and end are that range. Trim rejects a range longer than 900 seconds, which a quote never approaches. See the video trim docs for the exact request.
Two limits to remember: STT 1.0 has no speaker labels, and a quote the speaker garbled will not match, so also try a three- or four-word fragment.
Using the numbers
Once you have the start and end of a quote, you can do more than cut a clip. Build a table of contents from the first occurrence of each topic phrase, link a pull quote to the second it is spoken, or check that a speaker said the sentence a press release attributes to them. Keep the original recording and the full word list so any clip can be re-derived.
If the same phrase appears several times, collect every match instead of returning the first. Extend the function to yield each hit and present the list to a human. For recordings over 600 seconds, remember that each slice has its own clock, so add the slice's start offset before you report a time.
Sources
Related posts
More in Developers
- First-frame image for /v1/videos: public HTTPS only, no signed URLs
Image and video inputs to Sume generation must be fetchable public HTTPS URLs. Localhost, private IPs, signed URLs and wrong content types are rejected.
- Free, Pro, Startup, Scale: processing seats, queue slots, full hold
Sume's concurrency by plan, queue capacity max(3, 5 x concurrency), accepted job capacity, and the balance reserved if every slot holds a 10 s clip.
- Typed Sume video client from the OpenAPI JSON, after Sora
Sume publishes OpenAPI 3.0.3 at api.sume.com/reference/json. List the four video operations, then use the SDK's generated calls instead of hand-typing the wire.
- Generate then cut out: an image-model result into RMBG, in Python
Two Sume calls: generate a product shot, then POST its URL to /v1/rmbg-1.0/remove. Runnable Python, the polling loop, and the $0.1225 total per cutout.
Written by Sume