Find each script line in a recorded read: match STT words to script
Transcribe a studio read with Sume STT, then match its words to your script lines to get start and end seconds per line for cutting takes. Python.

A voice actor reads 40 lines into one 8-minute file. The game or video needs 40 clips. Cutting by ear takes an afternoon. Cutting from timestamps takes a minute, if you know where each script line starts and ends. Sume STT 1.0 gives word times for the recording, and the script gives the words to look for.
The matching idea
Normalize the script into lines of tokens. Walk through the transcribed words in order, and for each script line find the span that begins at the line's first token and covers about as many words as the line has. Because a read can have stumbles and repeats, anchor on the first three words and the last three words of each line. Take the last time the first anchor appears before the next line's start, so a retake wins over a flub.
Code
This version assumes the actor read the lines in order and does not repeat lines. Add retake handling if they do.
import re
def norm(s): return re.sub(r"[^a-z0-9']", "", s.lower())
def locate(words, lines):
toks = [(norm(w["word"]), w["start"], w["end"]) for w in words if w.get("type") == "word"]
out, pos = [], 0
for line in lines:
want = [norm(x) for x in line.split()]
head, tail = want[:3], want[-3:]
a = next((i for i in range(pos, len(toks) - len(head) + 1)
if [t[0] for t in toks[i:i + len(head)]] == head), None)
if a is None:
out.append((line, None)); continue
b = next((i for i in range(a, len(toks) - len(tail) + 1)
if [t[0] for t in toks[i:i + len(tail)]] == tail), None)
if b is None:
out.append((line, None)); continue
end = b + len(tail) - 1
out.append((line, (toks[a][1], toks[end][2])))
pos = end + 1
return outFrom spans to clips
Each hit gives a start and end in seconds. Pad them by 0.1 to 0.2 seconds at both ends. POST /v1/timeline-1.0/audio in split mode takes a url and 1 to 20 ranges of start and end, at $0.01 per job, for audio that is on media.sume.com (import it first with POST /v1/media-imports). The timeline audio docs have the details. Forty lines need two split jobs, so about 2 cents.
Edges and costs
- Lines that return
Nonewere not found. Listen to those yourself. The cause is usually a changed word, and sometimes a missed line. - One STT job takes at most 10 minutes, about 10 cents at $0.01 per minute. A longer read needs parts, with their offsets added to the word times.
- Names, numbers and invented words are where transcription differs most from your script. Anchor on plain words where you can.
- Always re-listen to the first and last cuts. A clean cut is the one you hear.
Sources
Related posts
More in Use cases
- Fit a voiceover to a 30-second ad: measure pace, then set TTS speed
Send timestamps.words, read the last word end time, and rescale generation_config.speed (0.6 to 1.5) until a Sume TTS take fits its slot. Two takes, 2 cents.
- Five product photos to one 15-second ad: five clips and a Timeline
Animate five angles of one product with Wan 3.0 at $0.625 each, cut 3 s from each in a $0.10 Timeline render, add $0.20 captions: $3.43 for a 15 s vertical ad.
- Fix one detail in a finished AI video: three routes on Sume
Seedance 2.5 mentions local editing. On Sume your options are Omni edit, Recast for a person swap, or a new take. Pick by what must stay unchanged.
- Fix one typo in an AI-generated image: quote the old and new text
Send the image as the first input_references entry and name the wrong text and the right text in quotes. Ideogram 4.5 on Sume costs $0.0375 to $0.275 per try.
Written by Sume