Search a podcast archive for a phrase and jump to the time: STT words

Transcribe episodes with Sume STT, keep words[] with start times in a small index, and find any phrase with a mm:ss link. Python, about 60 cents per hour.

5 min readSume
All posts

Someone remembers that a guest said "churn is a pricing problem" on one of 80 episodes, and nobody remembers which. A text search over full transcripts finds the episode, and a word time finds the second. Sume STT 1.0 returns both in words[]: the token, start and end in seconds, and a type of word or spacing.

Build the index once

Keep one row per word: episode id, offset, normalized token, start. If an episode runs past the 10-minute job limit, split it into 10-minute parts and add the part's offset to every word time. Searching is then a scan for a token sequence. For a few hundred hours, a list in memory or one SQLite table is enough; you do not need a search engine.

import re, sqlite3
db = sqlite3.connect("archive.db")
db.execute("create table if not exists w (ep text, i integer, tok text, t real)")

def norm(s):
    return re.sub(r"[^a-z0-9']", "", s.lower())

def add_episode(ep, words, offset=0.0):
    rows = [(ep, i, norm(w["word"]), w["start"] + offset)
            for i, w in enumerate(w for w in words if w.get("type") == "word")]
    db.executemany("insert into w values (?,?,?,?)", rows)
    db.commit()

def find(phrase):
    want = [norm(p) for p in phrase.split()]
    first = db.execute("select ep, i, t from w where tok = ?", (want[0],)).fetchall()
    for ep, i, t in first:
        got = [r[0] for r in db.execute(
            "select tok from w where ep = ? and i >= ? order by i limit ?", (ep, i, len(want)))]
        if got == want:
            yield ep, int(t // 60), int(t % 60)

Query it

Call find("churn is a pricing problem") and print each hit as episode 41 at 23:07. If you host the audio, add #t=1387 to the player link so listeners land on the right second. Make the query tolerant of the way people remember things. Search for a three-word core, not the full sentence.

Cost

STT 1.0 is about $0.01 per audio minute, or about 60 cents per hour. An archive of 80 episodes at 45 minutes is 3,600 minutes, or about $36, paid once. Searching afterwards costs nothing, since the index is yours. A 45-minute episode is five 10-minute jobs, so name each with an idempotency key of the form ep41-p03 and a rerun does not pay for parts twice.

Limits to be honest about

  • There are no speaker labels, so you can find what was said, not who said it.
  • The provider writes numbers and names its own way. If you search for "2026", also try "twenty twenty six".
  • Provider words can be wrong. A miss in the index does not prove the phrase was never spoken.
  • words[] is capped at 20,000 tokens per job, which a 10-minute job does not reach.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume