Search a podcast archive for a phrase and jump to the time: STT words
Transcribe episodes with Sume STT, keep words[] with start times in a small index, and find any phrase with a mm:ss link. Python, about 60 cents per hour.

Someone remembers that a guest said "churn is a pricing problem" on one of 80 episodes, and nobody remembers which. A text search over full transcripts finds the episode, and a word time finds the second. Sume STT 1.0 returns both in words[]: the token, start and end in seconds, and a type of word or spacing.
Build the index once
Keep one row per word: episode id, offset, normalized token, start. If an episode runs past the 10-minute job limit, split it into 10-minute parts and add the part's offset to every word time. Searching is then a scan for a token sequence. For a few hundred hours, a list in memory or one SQLite table is enough; you do not need a search engine.
import re, sqlite3
db = sqlite3.connect("archive.db")
db.execute("create table if not exists w (ep text, i integer, tok text, t real)")
def norm(s):
return re.sub(r"[^a-z0-9']", "", s.lower())
def add_episode(ep, words, offset=0.0):
rows = [(ep, i, norm(w["word"]), w["start"] + offset)
for i, w in enumerate(w for w in words if w.get("type") == "word")]
db.executemany("insert into w values (?,?,?,?)", rows)
db.commit()
def find(phrase):
want = [norm(p) for p in phrase.split()]
first = db.execute("select ep, i, t from w where tok = ?", (want[0],)).fetchall()
for ep, i, t in first:
got = [r[0] for r in db.execute(
"select tok from w where ep = ? and i >= ? order by i limit ?", (ep, i, len(want)))]
if got == want:
yield ep, int(t // 60), int(t % 60)Query it
Call find("churn is a pricing problem") and print each hit as episode 41 at 23:07. If you host the audio, add #t=1387 to the player link so listeners land on the right second. Make the query tolerant of the way people remember things. Search for a three-word core, not the full sentence.
Cost
STT 1.0 is about $0.01 per audio minute, or about 60 cents per hour. An archive of 80 episodes at 45 minutes is 3,600 minutes, or about $36, paid once. Searching afterwards costs nothing, since the index is yours. A 45-minute episode is five 10-minute jobs, so name each with an idempotency key of the form ep41-p03 and a rerun does not pay for parts twice.
Limits to be honest about
- There are no speaker labels, so you can find what was said, not who said it.
- The provider writes numbers and names its own way. If you search for "2026", also try "twenty twenty six".
- Provider words can be wrong. A miss in the index does not prove the phrase was never spoken.
words[]is capped at 20,000 tokens per job, which a 10-minute job does not reach.
Sources
Related posts
More in Use cases
- Swap the seasonal dish in a restaurant clip with Omni Flash edit
Keep one good restaurant clip and edit the dish for each holiday special with Gemini Omni Flash 1.1 video_to_video: $0.125 per output second at 720p.
- Seedance 2.5 takes 30 image references per pass: a shot-list ad
ByteDance lists 30 images, 10 clips and 10 audio files per Seedance 2.5 pass. How to turn a shot list into one seedance-2.5 request on Sume, with limits.
- Go past 30 s with Seedance 2.5 on Sume: hand off the last frame
Seedance 2.5 extends clips in rounds on the vendor side. On Sume, pull a late frame with video frames, feed it to the next job as a first frame, then join.
- Seedance 2.5 reference inventory: 30 images, 10 clips, 10 tracks
A pre-flight checklist for the 30 image, 10 video and 10 audio reference slots Seedance 2.5 offers, and how to check what Sume accepts.
Written by Sume