What is a critical error in TTS? A rubric after Nova 2 Sonic's 28%
Amazon reports 28% fewer critical errors in Nova 2 Sonic on an internal set and does not define them. Define yours, then tally with a read-back on Sume.

Amazon's release notes report 28% fewer critical errors in the May 2026 Nova 2 Sonic refresh, measured on an internal data set, and the page I read does not define the term. So write your own definition before you use anyone's figure. A workable one has three levels: critical means a wrong number, name or code is spoken; major means a wrong word that changes meaning; minor means a pacing or stress flaw. Then count them per 100 lines on your own scripts.
A three-level rubric
A rubric turns opinions into counts. It also tells you which failures to automate. Critical errors in the sense above can be caught by a machine read-back, because they are text differences. Minor errors need ears. Decide up front what blocks a release.
| Level | Example | Detect with | Action |
|---|---|---|---|
| Critical | Price, date or code spoken wrong | Read-back diff on must-keep strings | Block release; regenerate |
| Major | A word swapped, added or dropped | Read-back word diff | Fix text or voice, regenerate |
| Minor | Odd stress, flat pacing | Human listen on a sample | Tune speed; note it |
Tally it on Sume
Generate each line as its own job so you can see which line failed. A line that fails costs you a regenerate at 1 cent for up to 210 characters, or 2 cents up to 421 characters. Sample 5% of lines for human listening, and run the read-back on all of them. Keep the counts per model id and date, so a change in the rate is visible after a vendor refresh.
def severity(must_keep, transcript, ref, norm=lambda s: "".join(c for c in s.lower() if c.isalnum())):
t = norm(transcript)
if any(norm(m) not in t for m in must_keep):
return "critical"
if norm(ref) != t:
return "major"
return "ok"
print(severity(["$1,204.50"], "your total is 1204 50", "Your total is $1,204.50"))
print(severity(["4F7Q92"], "code four f seven q nine two", "Code 4F7Q92"))Why a rubric beats a percentage
A vendor's '28% fewer' is a ratio of two unknown counts. If it fell from 4 critical errors per thousand lines to 3, you would barely notice. If it fell from 40 to 29, you would. Your counts give you the base rate, and the base rate is what tells you how much a refresh can matter to you.
Keep the rubric stable across quarters, and write the date of every tally next to it.
Setting a release gate
Pick a gate in numbers: zero critical errors in the sample, and fewer than three major errors per hundred lines. When a line fails, regenerate only that line, and keep the failed job's id in your log. A pass applies to the exact model id and voice used, so a change to either restarts the tally. Treat the rubric as part of your release checklist, the same as a spell check.
Where the rubric came from
The three levels mirror how a listener reacts. A wrong price makes them act on false information. A swapped word makes them pause and re-read. A flat stress pattern barely registers. Weighting the levels 100, 10 and 1 gives a single score per hundred lines, but keep the raw counts as well, since a score can hide one critical error among many minor ones.
Sources
Related posts
More in Use cases
- What counts as original for an AI-made YouTube Short?
YouTube's Oct 2026 Shorts update favors original work. Here is what its own pages say about AI, templates and edits, and what they leave unsaid.
- What size should a TikTok video be? 2026 ad minimums by type
TikTok's own ad pages list 9:16 at 540x960 minimum for in-feed and 720x1280 for App Bundle. See the table and the Sume parameters that hit each one.
- Which AI video model for ads, product shots or talking heads?
On Sume, pick Wan 3.0 or Omni for ads, Kling 3 for silent product shots, and the Avatar Video route for talking heads. One table with rates and limits.
- Which LinkedIn ad formats can carry a Lead Gen Form?
LinkedIn lists single image, carousel, video, event, message, document and conversation ads for Lead Gen Forms. See which asset Sume can produce for each.
Written by Sume