Check an AI voiceover by transcribing it: a script diff for 4 cents
Run Sume STT on a finished TTS take and diff it against the script to catch misread numbers and names. 4 cents per 450-character line, TTS plus the check.

To catch a misread number or brand name in an AI voiceover without listening to every take, transcribe the finished audio with Sume STT and diff the transcript against the script. For a 450-character line that is 3 cents for the TTS take and 1 cent for the check, 4 cents in all; 100 lines come to $4.00 instead of $3.00.
The check finds clear errors such as a dropped sentence or "2026" read as "twenty twenty". It does not judge tone, pacing or accent, and a recognizer can mis-hear a correct take, so treat a flagged line as one to listen to, not as proof of a bad take.
What the check costs
Sume TTS 1.0 bills $0.0475 per 1,000 characters, rounded up per job, so 450 characters is 3 cents. STT bills about a cent per audio minute rounded up with a one-cent minimum, and a 30-second read is one cent. Both numbers are the published Sume rates; the speaking rate behind 450 characters as about 30 seconds is an assumption.
| Step | Route | Cost per line | 100 lines |
|---|---|---|---|
| TTS take | POST /v1/tts-1.0/generate | $0.03 | $3.00 |
| Check | POST /v1/stt-1.0/transcribe | $0.01 | $1.00 |
| Total | $0.04 | $4.00 |
The diff
Normalize both texts the same way, then use a sequence diff and print only the differing spans. Numbers are the common false alarm: a script written as "20%" will transcribe as "20 percent" or "twenty percent", so write the script the way you want it spoken for the check, or add a small replacement map.
import re, difflib
def toks(s):
return re.sub(r"[^\w\s']", " ", s.lower()).split()
def flag(script, heard):
a, b = toks(script), toks(heard)
sm = difflib.SequenceMatcher(None, a, b)
return [(op, " ".join(a[i1:i2]), " ".join(b[j1:j2]))
for op, i1, i2, j1, j2 in sm.get_opcodes() if op != "equal"]
print(flag("Use code POD for 20 percent off",
"Use code pod for twenty percent off"))When it is worth it
Use it on lines with numbers, prices, dates, URLs and product names, where one wrong word costs money. Skip it on filler lines. If a take fails, a retake is another 3 cents, and the brand-name post lists four levers to fix pronunciation first, including the optional pronunciation_dict_id on the TTS request.
Limits of the method
A transcript check proves that words were spoken, not that they were spoken well. It cannot hear a wrong stress, a flat read or a foreign accent on a name. It can also flag a correct take if the recognizer mishears a rare word, which is why the output of flag is a list to review, not an automatic reject.
Use the check as a gate in a batch: generate all takes, transcribe them, diff, and send only the flagged lines to a human. On a 100-line batch where five lines are flagged, a person listens to five clips instead of a hundred.
Sources
Related posts
More in Media tools
- Trim and conform to 1080x1920 at 30 fps in one video-trim job
Video trim's output field re-encodes to a set width, height and 24/25/30/60 fps in the same $0.02 job, exact precision only. Request, ranges and refusal codes.
- Crop a 21:9 or 2.39:1 film clip to 9:16 with video filter
Width fractions for cropping ultrawide 21:9, 2.39:1 and 32:9 clips to 9:16 with one Sume video-filter crop op, plus the 0.05 minimum side and a free check.
- Trim before crop: how the 300 s filter limit sets the order
Video filter accepts sources up to 300 seconds while video trim accepts 1800. For a 20-minute clip, trim first, then crop: two jobs, $0.04.
- Dim a background clip under text: video filter dim amounts
The Sume video-filter dim op multiplies luma by an amount between 0 and 1; black stays black and chroma is untouched. Amount table, request and the $0.02 price.
Written by Sume