Check a voiceover with STT: 312 seconds costs 5.2 cents

Transcribe the finished voiceover and diff the words against the script: Sume STT bills $0.052 for 312 s. A runnable Python word-diff included. Read 2026-10-08.

5 min readSume
All posts

A 312-second voiceover costs $0.052 to transcribe on Sume STT (5.2 minutes x $0.01), and a word-level diff of that transcript against the script shows dropped, swapped or doubled words before you render. The same audio is $0.0087 on MAI-Transcribe-2 batch at Microsoft's $0.10 an hour. Neither check is free of cost, but both are cheaper than one extra render.

Cost of the check

The cost is linear in minutes. Sume's request cap is 600 seconds, so a 312-second voiceover is one request; a 1,000-second one is two. Microsoft's page prices are introductory.

Cost of transcribing a finished voiceover, read 2026-10-08
Audio lengthSume STTMAI-Transcribe-2 batchSume requests
60 s$0.01$0.001671
312 s$0.052$0.008671
600 s$0.10$0.016671
1,000 s$0.1667$0.027782

The diff

This script runs as written and uses only the standard library. It lowercases both texts, splits into words, and prints every place where the transcript differs from the script. The STT result carries words[] with times, so in a real run you would take the words from the job result instead of a string.

import difflib, re
script = "Two pears for three dollars today only at the corner market"
heard = "Two pairs for three dollars today at the corner market"
norm = lambda s: re.findall(r"[a-z0-9']+", s.lower())
a, b = norm(script), norm(heard)
for op, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
    if op != "equal":
        print(op, a[i1:i2], "->", b[j1:j2])
print("script words", len(a), "heard words", len(b))

What the diff can and cannot show

A word-level diff catches dropped and replaced words, which are the most common faults a listener notices. It does not judge pronunciation or pacing, and the recognizer can make its own mistakes (here it hears pairs for pears), so treat each flagged difference as a prompt to listen at that timestamp. Because the transcript has word times, you can jump straight to the second where the mismatch begins. For a Timeline build, run the check on the joined spine, not on each line.

Where the check sits in a pipeline

Run the check between speech and render. A Timeline render is billed per started minute, so a voiceover that is wrong costs the render plus the redo; a check that costs 5.2 cents on a 312-second read costs less than the 10 cents of a single render minute. For a script that goes through several TTS jobs, joined with Timeline audio, transcribe the joined file, since a missing word at a seam is exactly what a per-line check misses.

If the diff finds a problem, the repair is local: regenerate the one line, join again at $0.01, and transcribe again. For a 312-second spine the whole loop is about 5.2 cents plus one TTS job plus $0.01, and the loop can repeat until the diff is clean.

A tolerance helps. Recognizers drop punctuation and sometimes normalize numbers, so compare words after lowercasing and, if your script has numerals, write them out as the recognizer would. The script above ignores punctuation, which keeps the output short. Add a small allow-list of known variants, such as a brand name spelled two ways, so the report shows only real differences.

Cost over a month

At 40 voiceovers a month averaging 312 seconds, the Sume check costs 40 x $0.052 = $2.08, and the Microsoft batch price (an introductory $0.10 an hour) is 40 x $0.00867 = $0.35. The difference is $1.73 a month, which is the price of keeping the transcript inside the same job system.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume