Find the timestamp of a quote in a recording with Sume STT words

Sume's STT result returns every word with start and end seconds. A short Python function finds a quoted phrase and returns where to cut. About a cent a minute.

5 min readSume
All posts

To find where a quote is spoken, transcribe the file with Sume STT 1.0 and search the returned words array, each entry {word, start, end} in seconds from the audio start. A match of the quote's words gives you the start of the first word and the end of the last. A 10-minute file costs 10 cents.

The function below is plain Python with no dependencies, and it was run against a small sample before publishing.

What the result contains

The result of a completed STT job has text, language_code, language_probability and words. Words are ordered by start, and the array is always present, even when empty. It is capped at 20,000 entries, which a 600-second transcript stays well under. Some tokens may be classed spacing in the optional type field, so skip those when matching.

STT 1.0 takes audio up to 600 seconds. For a longer recording, split it and add the slice offset to each match.

Sume STT 1.0 result fields used for quote search (read 2026-10-08)
FieldMeaningUnit
words[].wordToken textstring
words[].startToken start from audio startseconds
words[].endToken endseconds
words[].typeProvider class, for example spacingoptional
words_truncatedPresent only if capped at 20,000boolean

The search function

Normalize both sides by lowercasing and stripping punctuation, then slide a window across the word list. Pad the cut by a fraction of a second so the clip does not clip the first syllable.

import re

def norm(s):
    return re.sub(r"[^\w']+", "", s.lower())

def find_quote(words, quote, pad=0.2):
    toks = [w for w in words if w.get("type") != "spacing"]
    want = [norm(t) for t in quote.split() if norm(t)]
    have = [norm(w["word"]) for w in toks]
    for i in range(len(have) - len(want) + 1):
        if have[i:i + len(want)] == want:
            return (max(0.0, toks[i]["start"] - pad),
                    toks[i + len(want) - 1]["end"] + pad)
    return None

sample = [{"word": "We", "start": 1.0, "end": 1.2},
          {"word": "ship", "start": 1.2, "end": 1.6},
          {"word": "on", "start": 1.6, "end": 1.7},
          {"word": "Friday.", "start": 1.7, "end": 2.3}]
print(find_quote(sample, "ship on friday"))

Then cut the clip

If the recording is a video you uploaded to Sume's media host, the trim route cuts a [start, end) range in seconds, and the quote's start and end are that range. Trim rejects a range longer than 900 seconds, which a quote never approaches. See the video trim docs for the exact request.

Two limits to remember: STT 1.0 has no speaker labels, and a quote the speaker garbled will not match, so also try a three- or four-word fragment.

Using the numbers

Once you have the start and end of a quote, you can do more than cut a clip. Build a table of contents from the first occurrence of each topic phrase, link a pull quote to the second it is spoken, or check that a speaker said the sentence a press release attributes to them. Keep the original recording and the full word list so any clip can be re-derived.

If the same phrase appears several times, collect every match instead of returning the first. Extend the function to yield each hit and present the list to a human. For recordings over 600 seconds, remember that each slice has its own clock, so add the slice's start offset before you report a time.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume