Nova 2 Sonic cut hallucinations 88%: run a read-back check on Sume
AWS says its May Nova 2 Sonic refresh cut hallucinations 88% on an internal set. Test any voice on your own text: Sume TTS, then STT, then a diff.

Do not take a vendor's hallucination number as your number. Amazon's own release notes put the May 2026 Nova 2 Sonic refresh at 88% fewer hallucinations, and in the same sentence say the figure is measured on an internal data set. The only figure that matters for you is the one from your scripts, so build a read-back check: synthesize the text, transcribe the audio, and diff the two. On Sume that is two jobs and about a dozen lines of Python.
Source: Amazon Nova 2 release notes (read 2026-10-05).
What the release notes actually claim
The May 2026 Nova 2 Sonic refresh lists three reductions: hallucinations down 88%, speaker drift down 52%, and critical errors down 28%. The page labels all three as measured on an internal data set. It also names verbatim fidelity for alphanumeric codes, currency, addresses, emails and phone numbers as an improvement area. That is a useful list of things to test, because those are the strings a listener cannot easily forgive getting wrong.
| Claim | Figure | Basis stated on the page |
|---|---|---|
| Hallucinations | -88% | Internal data set |
| Speaker drift | -52% | Internal data set |
| Critical errors | -28% | Internal data set |
| Verbatim fidelity | Codes, currency, addresses, emails, phone numbers | Named as improved; no figure |
A read-back check on Sume
A hallucination in speech is audio that says words the text does not contain. A read-back check catches it without listening. Submit the text to the Sume TTS router, pass the result audio URL to Sume STT, normalise both strings and compare. Sume TTS is billed at $0.0475 per 1,000 characters (ceil to cents, 1 cent minimum) and STT at $0.01 per audio minute, so a 200-character line costs 1 cent to speak, and a 10-second clip costs a fraction of a cent to transcribe at the minute rate.
import os, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
def run(path, body):
d = requests.post(B + path, json=body, headers=H, timeout=60).json()["data"]
while not d.get("terminal"):
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
return requests.get(d["result_url"], headers=H, timeout=60).json()
text = "Order 4F7Q-92 ships to 18 Elm St for $1,204.50."
tts = run("/v1/tts-router/generate", {"model": "sonic-3.6", "transcript": text,
"voice": {"id": os.environ["SUME_VOICE_ID"]}})
print(tts) # find the audio URL in the result, then send it to STT:
# stt = run("/v1/stt-1.0/transcribe", {"audio_url": AUDIO_URL, "duration_seconds": 10})What to compare and what to ignore
Strip case and punctuation, then compare alphanumerics only. A transcriber may write 'twelve hundred' for 1,200, so keep a list of the strings that must survive exactly: order codes, prices, emails. Count a run as failed when any must-survive string is missing from the transcript. This tests the whole chain, so an STT miss shows up as a false alarm; re-run the failures once before you blame the voice.
A passing read-back proves the audio contained those words. It does not prove the text you submitted was the approved text. For that, use the source-bound receipt on Sume TTS jobs: the receipt holds the SHA-256 of the submitted transcript, and it proves the submitted text, not the pronunciation. The two checks cover different failures, so run both.
Cost of running it on a real script
Check 50 lines of 200 characters each. Each line is 200 x 0.00475 = 0.95 cents, which rounds up to 1 cent, so TTS is 50 cents. If each line is about 10 seconds long, STT reads 50 x 10 s = 500 s of audio, about 8.3 minutes, roughly 8 cents at the published minute rate. The whole check is under a dollar, which is cheap against shipping a wrong price aloud.
Sources
Related posts
More in Developers
- 503 api_key_auth_unavailable on Sume: your key is fine
A 503 api_key_auth_unavailable or api_key_auth_not_configured is Sume's auth check failing, not a bad key. Keep the key, back off, and read request_id.
- AppleScript: submit a Wan 3.0 video job and save it to the Desktop
A 22-line AppleScript that runs curl and jq through do shell script, polls a Sume video job and saves the MP4 to the Desktop. A 2 s 480p Wan 3.0 clip is $0.125.
- Ask a decision model to approve a Sume dry_run estimate before paying
Call the paid MCP tool with dry_run true, hand the estimate to a yes/no decision, and only then send idempotency_key with a max_spend_usd that you set.
- Entity error 14.41%? Score your own call audio with Sume STT in Python
AssemblyAI reports 14.41% entity error on voice-agent audio and 3.44% English WER. Neither is yours. Compute entity recall on 20 of your clips with Sume STT.
Written by Sume