Check a voiceover with STT: 312 seconds costs 5.2 cents
Transcribe the finished voiceover and diff the words against the script: Sume STT bills $0.052 for 312 s. A runnable Python word-diff included. Read 2026-10-08.

A 312-second voiceover costs $0.052 to transcribe on Sume STT (5.2 minutes x $0.01), and a word-level diff of that transcript against the script shows dropped, swapped or doubled words before you render. The same audio is $0.0087 on MAI-Transcribe-2 batch at Microsoft's $0.10 an hour. Neither check is free of cost, but both are cheaper than one extra render.
Cost of the check
The cost is linear in minutes. Sume's request cap is 600 seconds, so a 312-second voiceover is one request; a 1,000-second one is two. Microsoft's page prices are introductory.
| Audio length | Sume STT | MAI-Transcribe-2 batch | Sume requests |
|---|---|---|---|
| 60 s | $0.01 | $0.00167 | 1 |
| 312 s | $0.052 | $0.00867 | 1 |
| 600 s | $0.10 | $0.01667 | 1 |
| 1,000 s | $0.1667 | $0.02778 | 2 |
The diff
This script runs as written and uses only the standard library. It lowercases both texts, splits into words, and prints every place where the transcript differs from the script. The STT result carries words[] with times, so in a real run you would take the words from the job result instead of a string.
import difflib, re
script = "Two pears for three dollars today only at the corner market"
heard = "Two pairs for three dollars today at the corner market"
norm = lambda s: re.findall(r"[a-z0-9']+", s.lower())
a, b = norm(script), norm(heard)
for op, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
if op != "equal":
print(op, a[i1:i2], "->", b[j1:j2])
print("script words", len(a), "heard words", len(b))What the diff can and cannot show
A word-level diff catches dropped and replaced words, which are the most common faults a listener notices. It does not judge pronunciation or pacing, and the recognizer can make its own mistakes (here it hears pairs for pears), so treat each flagged difference as a prompt to listen at that timestamp. Because the transcript has word times, you can jump straight to the second where the mismatch begins. For a Timeline build, run the check on the joined spine, not on each line.
Where the check sits in a pipeline
Run the check between speech and render. A Timeline render is billed per started minute, so a voiceover that is wrong costs the render plus the redo; a check that costs 5.2 cents on a 312-second read costs less than the 10 cents of a single render minute. For a script that goes through several TTS jobs, joined with Timeline audio, transcribe the joined file, since a missing word at a seam is exactly what a per-line check misses.
If the diff finds a problem, the repair is local: regenerate the one line, join again at $0.01, and transcribe again. For a 312-second spine the whole loop is about 5.2 cents plus one TTS job plus $0.01, and the loop can repeat until the diff is clean.
A tolerance helps. Recognizers drop punctuation and sometimes normalize numbers, so compare words after lowercasing and, if your script has numerals, write them out as the recognizer would. The script above ignores punctuation, which keeps the output short. Add a small allow-list of known variants, such as a brand name spelled two ways, so the report shows only real differences.
Cost over a month
At 40 voiceovers a month averaging 312 seconds, the Sume check costs 40 x $0.052 = $2.08, and the Microsoft batch price (an introductory $0.10 an hour) is 40 x $0.00867 = $0.35. The difference is $1.73 a month, which is the price of keeping the transcript inside the same job system.
Sources
Related posts
More in Developers
- Claude Code hook: exit 2 or a JSON deny for a paid Sume call?
Exit code 2 blocks and cannot be overridden by JSON; exit 0 with permissionDecision deny carries a reason. A runnable Python guard for Sume's paid tools.
- Hook matcher rules: when mcp__sume__ names are exact, not regex
Claude Code reads a matcher of letters, digits, underscores and pipes as exact names, and other characters as regex. How to write it for Sume tools.
- Claude Code http hook: send paid Sume call records to an audit log
An http hook POSTs hook input to a URL, with headers from allowed environment variables only. Log Sume tool names and idempotency keys with a tiny receiver.
- Claude per-message effort: drop to low while a Sume job runs
Claude's per-message effort beta keeps the prompt cache when you change effort mid-conversation. How to use it around Sume jobs, and its Haiku 5.5 limit.
Written by Sume