Check a TTS take with an STT round trip: flag skipped words in Python
Transcribe your own voiceover with Sume STT and diff it against the script to catch skipped or changed words before you ship. Python sketch included.

Yes: send the finished voiceover to Sume STT, normalize both the script and the transcript to lowercase words, and diff them. Any word the transcript lacks is a spot to listen to. It takes one STT job per take and a few lines of Python.
The check flags where to listen. It does not prove the audio is right, because speech-to-text makes its own mistakes.
Why re-check a take after a model or voice change?
New speech models keep shipping. TechCrunch reported ElevenLabs v4 on 2026-09-28 with 90+ languages, and Google's Gemini API changelog lists Gemini 3.8 Flash TTS as generally available on 2026-09-22. Sume TTS 1.0 has no engine picker, so you do not switch models by name there, but voices and scripts still change. A cheap regression check beats finding a dropped word after publishing.
Numbers, product names and long sentences are where a dropped or changed word is easiest to miss by ear.
How do I run the round trip on Sume?
Create the TTS job and keep the audio artifact URL from result.artifacts[]. Then call POST /v1/stt-1.0/transcribe with audio_url set to that URL. The docs say to prefer a Sume media URL, and non-English audio should carry language_code (omit it for auto-detect). Poll both jobs by status_url; never resubmit a paid job just because your poll timed out.
Keep each take under 10 minutes. The STT duration_seconds field takes 1 to 600 and is used to reserve usage.
| Step | Sume field | Use |
|---|---|---|
| Script | transcript | The text you sent to TTS |
| Take | result.artifacts[].url | audio_url for STT |
| Transcript | result.text | Words the STT job heard |
| Timings | result.words[] | Where in the take a gap sits |
What does the diff look like?
The sketch below compares a script with its transcript and prints the words that differ. It runs as written; paste in the real strings.
import difflib, re
script = "Our new bottle keeps drinks cold for 24 hours. Order today and get free shipping."
transcript = "Our new bottle keeps drinks cold for 24 hours. Order today and get shipping."
def words(text):
return re.findall(r"[a-z0-9']+", text.lower())
a, b = words(script), words(transcript)
for op, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
if op != "equal":
print(op, "script:", a[i1:i2], "heard:", b[j1:j2])
print("match ratio", round(difflib.SequenceMatcher(None, a, b).ratio(), 3))
What should I do with a mismatch?
Listen at that time, using the word timings from the STT result to jump there. If the take really dropped the word, retake only that sentence and splice it in (see the related posts). If the transcript simply misheard a brand name, note it and move on. A match ratio is a triage number, so set your own bar and read the flagged words yourself.
Sources
Related posts
More in Developers
- Check a video request against /v1/videos/models before you submit
Duration, resolution, size and seed errors cost a round trip. A short Python validator reads the model catalog and refuses a bad request locally first.
- Check a transparent GPT Image 2.5 PNG for real alpha in Python
A transparent GPT Image 2.5 result can still look opaque. Ask for background transparent as PNG, then check the alpha channel in Python: a 20-line script.
- Check avatar lip sync: transcript word times, then stills
Transcribe an avatar video with Sume Video inspect, take word timestamps from words[], and pull stills at those instants with video_frames to inspect the mouth.
- Check caption reading speed from STT segments and flag dense sentences
Request sentence segments from Sume STT, compute characters per second for each, and flag the ones too dense to read on screen. Python and caption fixes.
Written by Sume