Check an AI voiceover by transcribing it: a script diff for 4 cents

Run Sume STT on a finished TTS take and diff it against the script to catch misread numbers and names. 4 cents per 450-character line, TTS plus the check.

5 min readSume
All posts

To catch a misread number or brand name in an AI voiceover without listening to every take, transcribe the finished audio with Sume STT and diff the transcript against the script. For a 450-character line that is 3 cents for the TTS take and 1 cent for the check, 4 cents in all; 100 lines come to $4.00 instead of $3.00.

The check finds clear errors such as a dropped sentence or "2026" read as "twenty twenty". It does not judge tone, pacing or accent, and a recognizer can mis-hear a correct take, so treat a flagged line as one to listen to, not as proof of a bad take.

What the check costs

Sume TTS 1.0 bills $0.0475 per 1,000 characters, rounded up per job, so 450 characters is 3 cents. STT bills about a cent per audio minute rounded up with a one-cent minimum, and a 30-second read is one cent. Both numbers are the published Sume rates; the speaking rate behind 450 characters as about 30 seconds is an assumption.

Per-line cost with and without the transcript check, 450 characters (read 2026-10-08)
StepRouteCost per line100 lines
TTS takePOST /v1/tts-1.0/generate$0.03$3.00
CheckPOST /v1/stt-1.0/transcribe$0.01$1.00
Total$0.04$4.00

The diff

Normalize both texts the same way, then use a sequence diff and print only the differing spans. Numbers are the common false alarm: a script written as "20%" will transcribe as "20 percent" or "twenty percent", so write the script the way you want it spoken for the check, or add a small replacement map.

import re, difflib

def toks(s):
    return re.sub(r"[^\w\s']", " ", s.lower()).split()

def flag(script, heard):
    a, b = toks(script), toks(heard)
    sm = difflib.SequenceMatcher(None, a, b)
    return [(op, " ".join(a[i1:i2]), " ".join(b[j1:j2]))
            for op, i1, i2, j1, j2 in sm.get_opcodes() if op != "equal"]

print(flag("Use code POD for 20 percent off",
           "Use code pod for twenty percent off"))

When it is worth it

Use it on lines with numbers, prices, dates, URLs and product names, where one wrong word costs money. Skip it on filler lines. If a take fails, a retake is another 3 cents, and the brand-name post lists four levers to fix pronunciation first, including the optional pronunciation_dict_id on the TTS request.

Limits of the method

A transcript check proves that words were spoken, not that they were spoken well. It cannot hear a wrong stress, a flat read or a foreign accent on a name. It can also flag a correct take if the recognizer mishears a rare word, which is why the output of flag is a list to review, not an automatic reject.

Use the check as a gate in a batch: generate all takes, transcribe them, diff, and send only the flagged lines to a human. On a 100-line batch where five lines are flagged, a person listens to five clips instead of a hundred.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume