What is a critical error in TTS? A rubric after Nova 2 Sonic's 28%

Amazon reports 28% fewer critical errors in Nova 2 Sonic on an internal set and does not define them. Define yours, then tally with a read-back on Sume.

5 min readSume
All posts

Amazon's release notes report 28% fewer critical errors in the May 2026 Nova 2 Sonic refresh, measured on an internal data set, and the page I read does not define the term. So write your own definition before you use anyone's figure. A workable one has three levels: critical means a wrong number, name or code is spoken; major means a wrong word that changes meaning; minor means a pacing or stress flaw. Then count them per 100 lines on your own scripts.

A three-level rubric

A rubric turns opinions into counts. It also tells you which failures to automate. Critical errors in the sense above can be caught by a machine read-back, because they are text differences. Minor errors need ears. Decide up front what blocks a release.

A TTS severity rubric and how to detect each level, with the Amazon figures it responds to, read 2026-10-05
LevelExampleDetect withAction
CriticalPrice, date or code spoken wrongRead-back diff on must-keep stringsBlock release; regenerate
MajorA word swapped, added or droppedRead-back word diffFix text or voice, regenerate
MinorOdd stress, flat pacingHuman listen on a sampleTune speed; note it

Tally it on Sume

Generate each line as its own job so you can see which line failed. A line that fails costs you a regenerate at 1 cent for up to 210 characters, or 2 cents up to 421 characters. Sample 5% of lines for human listening, and run the read-back on all of them. Keep the counts per model id and date, so a change in the rate is visible after a vendor refresh.

def severity(must_keep, transcript, ref, norm=lambda s: "".join(c for c in s.lower() if c.isalnum())):
    t = norm(transcript)
    if any(norm(m) not in t for m in must_keep):
        return "critical"
    if norm(ref) != t:
        return "major"
    return "ok"

print(severity(["$1,204.50"], "your total is 1204 50", "Your total is $1,204.50"))
print(severity(["4F7Q92"], "code four f seven q nine two", "Code 4F7Q92"))

Why a rubric beats a percentage

A vendor's '28% fewer' is a ratio of two unknown counts. If it fell from 4 critical errors per thousand lines to 3, you would barely notice. If it fell from 40 to 29, you would. Your counts give you the base rate, and the base rate is what tells you how much a refresh can matter to you.

Keep the rubric stable across quarters, and write the date of every tally next to it.

Setting a release gate

Pick a gate in numbers: zero critical errors in the sample, and fewer than three major errors per hundred lines. When a line fails, regenerate only that line, and keep the failed job's id in your log. A pass applies to the exact model id and voice used, so a change to either restarts the tally. Treat the rubric as part of your release checklist, the same as a spell check.

Where the rubric came from

The three levels mirror how a listener reacts. A wrong price makes them act on false information. A swapped word makes them pause and re-read. A flat stress pattern barely registers. Weighting the levels 100, 10 and 1 gives a single score per hundred lines, but keep the raw counts as well, since a score can hide one critical error among many minor ones.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume