Check a TTS take with an STT round trip: flag skipped words in Python

Transcribe your own voiceover with Sume STT and diff it against the script to catch skipped or changed words before you ship. Python sketch included.

4 min readSume
All posts

Yes: send the finished voiceover to Sume STT, normalize both the script and the transcript to lowercase words, and diff them. Any word the transcript lacks is a spot to listen to. It takes one STT job per take and a few lines of Python.

The check flags where to listen. It does not prove the audio is right, because speech-to-text makes its own mistakes.

Why re-check a take after a model or voice change?

New speech models keep shipping. TechCrunch reported ElevenLabs v4 on 2026-09-28 with 90+ languages, and Google's Gemini API changelog lists Gemini 3.8 Flash TTS as generally available on 2026-09-22. Sume TTS 1.0 has no engine picker, so you do not switch models by name there, but voices and scripts still change. A cheap regression check beats finding a dropped word after publishing.

Numbers, product names and long sentences are where a dropped or changed word is easiest to miss by ear.

How do I run the round trip on Sume?

Create the TTS job and keep the audio artifact URL from result.artifacts[]. Then call POST /v1/stt-1.0/transcribe with audio_url set to that URL. The docs say to prefer a Sume media URL, and non-English audio should carry language_code (omit it for auto-detect). Poll both jobs by status_url; never resubmit a paid job just because your poll timed out.

Keep each take under 10 minutes. The STT duration_seconds field takes 1 to 600 and is used to reserve usage.

What to compare in the round trip (Sume OpenAPI and docs, read 2026-10-02).
StepSume fieldUse
ScripttranscriptThe text you sent to TTS
Takeresult.artifacts[].urlaudio_url for STT
Transcriptresult.textWords the STT job heard
Timingsresult.words[]Where in the take a gap sits

What does the diff look like?

The sketch below compares a script with its transcript and prints the words that differ. It runs as written; paste in the real strings.

import difflib, re

script = "Our new bottle keeps drinks cold for 24 hours. Order today and get free shipping."
transcript = "Our new bottle keeps drinks cold for 24 hours. Order today and get shipping."

def words(text):
    return re.findall(r"[a-z0-9']+", text.lower())

a, b = words(script), words(transcript)
for op, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
    if op != "equal":
        print(op, "script:", a[i1:i2], "heard:", b[j1:j2])
print("match ratio", round(difflib.SequenceMatcher(None, a, b).ratio(), 3))

What should I do with a mismatch?

Listen at that time, using the word timings from the STT result to jump there. If the take really dropped the word, retake only that sentence and splice it in (see the related posts). If the transcript simply misheard a brand name, note it and move on. A match ratio is a triage number, so set your own bar and read the flagged words yourself.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume