Search what was said across clips: SQLite FTS5 from STT segments

Request sentence segmentation on Sume STT, load each segment into a SQLite FTS5 table, and find the clip and second where a phrase was spoken.

5 min readSume
All posts

To find which clip says a phrase and when, ask Sume STT 1.0 for sentence segments (segmentation: { mode: "sentence" }), insert each segment's text and start into a SQLite FTS5 table, and query it with MATCH. FTS5 ships with Python's standard sqlite3 on most builds, so there is nothing to host. Segment fields (index, text, start, end, duration_seconds) are in the STT result shape in the API reference, read 2026-10-06.

Ask for segments, not just text

Add segmentation to the submit body. Sume splits on terminal punctuation, falls back to a silence gap for text without punctuation, and joins words with single spaces. boundary_lead_ms (0 to 500, default 70) moves each segment start slightly earlier so a player lands just before the first word.

Fields in each STT segment, from the Sume API reference (docs.sume.com), read 2026-10-06.
FieldMeaning
indexPosition in the transcript
textThe sentence
start, endSeconds from the start of the audio
duration_secondsend minus start

Index and query

Keep the clip name in its own column so a hit tells you the file. The query printed both clips for shipping, with the match in brackets.

import sqlite3

db = sqlite3.connect(":memory:")
db.execute("CREATE VIRTUAL TABLE spoken USING fts5(clip, text, starts UNINDEXED)")


def add(clip, segments):
    for s in segments:
        db.execute("INSERT INTO spoken VALUES (?, ?, ?)", (clip, s["text"], s["start"]))


def find(query):
    return db.execute(
        "SELECT clip, starts, snippet(spoken, 1, '[', ']', '...', 8) "
        "FROM spoken WHERE spoken MATCH ? ORDER BY rank",
        (query,),
    ).fetchall()


add("a.wav", [{"text": "Free shipping on every order.", "start": 1.2}])
add("b.wav", [{"text": "Ask about shipping times.", "start": 7.9}])
print(find("shipping"))

Keeping the index tied to your files

Use your own file id as the clip name. Put it in the job's metadata or in the idempotency key and read it back from the result, as in matching results to file ids. To jump to the moment, convert starts to a frame with the 29.97 fps helper.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume