Find each script line in a recorded read: match STT words to script

Transcribe a studio read with Sume STT, then match its words to your script lines to get start and end seconds per line for cutting takes. Python.

5 min readSume
All posts

A voice actor reads 40 lines into one 8-minute file. The game or video needs 40 clips. Cutting by ear takes an afternoon. Cutting from timestamps takes a minute, if you know where each script line starts and ends. Sume STT 1.0 gives word times for the recording, and the script gives the words to look for.

The matching idea

Normalize the script into lines of tokens. Walk through the transcribed words in order, and for each script line find the span that begins at the line's first token and covers about as many words as the line has. Because a read can have stumbles and repeats, anchor on the first three words and the last three words of each line. Take the last time the first anchor appears before the next line's start, so a retake wins over a flub.

Code

This version assumes the actor read the lines in order and does not repeat lines. Add retake handling if they do.

import re
def norm(s): return re.sub(r"[^a-z0-9']", "", s.lower())

def locate(words, lines):
    toks = [(norm(w["word"]), w["start"], w["end"]) for w in words if w.get("type") == "word"]
    out, pos = [], 0
    for line in lines:
        want = [norm(x) for x in line.split()]
        head, tail = want[:3], want[-3:]
        a = next((i for i in range(pos, len(toks) - len(head) + 1)
                  if [t[0] for t in toks[i:i + len(head)]] == head), None)
        if a is None:
            out.append((line, None)); continue
        b = next((i for i in range(a, len(toks) - len(tail) + 1)
                  if [t[0] for t in toks[i:i + len(tail)]] == tail), None)
        if b is None:
            out.append((line, None)); continue
        end = b + len(tail) - 1
        out.append((line, (toks[a][1], toks[end][2])))
        pos = end + 1
    return out

From spans to clips

Each hit gives a start and end in seconds. Pad them by 0.1 to 0.2 seconds at both ends. POST /v1/timeline-1.0/audio in split mode takes a url and 1 to 20 ranges of start and end, at $0.01 per job, for audio that is on media.sume.com (import it first with POST /v1/media-imports). The timeline audio docs have the details. Forty lines need two split jobs, so about 2 cents.

Edges and costs

  • Lines that return None were not found. Listen to those yourself. The cause is usually a changed word, and sometimes a missed line.
  • One STT job takes at most 10 minutes, about 10 cents at $0.01 per minute. A longer read needs parts, with their offsets added to the word times.
  • Names, numbers and invented words are where transcription differs most from your script. Anchor on plain words where you can.
  • Always re-listen to the first and last cuts. A clean cut is the one you hear.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume