Make an AI voice read an email or order code right: spell, verify

Amazon and Cartesia both claim better codes, emails and phone numbers. The reliable fix is text prep plus a read-back check. Try both on Sume's TTS router.

5 min readSume
All posts

To make a synthetic voice read an email address or an order code correctly, spell the risky parts out in the text you submit, then verify the audio by transcribing it. Vendors are improving here: Amazon's May 2026 Nova 2 Sonic notes name verbatim fidelity for alphanumeric codes, currency, addresses, emails and phone numbers, and Cartesia's Sonic 3.6 page says confirmation codes and heteronyms are handled without preprocessing. A claim that you need less preprocessing is a reason to test it, not to skip the test.

Why codes and addresses fail

A listener hears a stream of sounds and cannot re-read it. A string like 4F7Q-92 is ambiguous: a model might say 'four F seven Q ninety-two' or 'four F seven Q nine two'. An email such as ana.lee@example.com has punctuation that must be spoken as words. If the text is ambiguous, the voice has to guess, so the first fix is to remove the ambiguity in what you send.

Claims about codes and emails, Amazon release notes and Cartesia Sonic page, read 2026-10-05
Vendor pageWhat it saysWhat you still need
Amazon Nova 2 release notesVerbatim fidelity for alphanumeric codes, currency, addresses, emails, phone numbersA check on your own strings
Cartesia Sonic 3.6Confirmation codes and heteronyms handled without preprocessingA check on your own strings
Sume TTSSubmitted text is stored; source-bound jobs carry a SHA-256 receiptA read-back to check pronunciation

Prepare the text, then submit

Write the string the way you want it heard: 'ana dot lee at example dot com' and 'four F seven Q, nine two'. Keep a function that does this for the formats you use, so every job gets the same treatment. Sume counts every character you send, including spaces and punctuation, so spelled-out text costs a little more: the 30-character line 'ana dot lee at example dot com' is 30 x 0.00475 = 0.1425 cents, which still bills as 1 cent.

import re

def speak_email(addr):
    return addr.replace(".", " dot ").replace("@", " at ").replace("-", " dash ")

def speak_code(code):
    return ", ".join(" ".join(part) for part in code.split("-"))

print(speak_email("ana.lee@example.com"))
print(speak_code("4F7Q-92"))
# -> ana dot lee at example dot com
# -> 4 F 7 Q, 9 2

Verify by listening with a machine

Submit the prepared line to POST /v1/tts-router/generate, send the audio to POST /v1/stt-1.0/transcribe, and check that the key strings come back, as in the read-back post. A receipt proves what text you submitted. It cannot prove the voice said it well, because the Sume docs state that a valid receipt proves the submitted text, not the pronunciation. So keep both: the receipt for text, the read-back for sound.

Run the check on a list of 20 real strings once, then again whenever a vendor announces a refresh.

Phone numbers and currency

The same approach covers the other strings on Amazon's list. Group a phone number into the chunks you want spoken, such as 'five five five, oh one two three', and write currency as words: 'one thousand two hundred four dollars and fifty cents' instead of '$1,204.50'. Longer text costs a little more, so a 60-character spelled price is 60 x 0.00475 = 0.285 cents, still 1 cent, and the price of a mistaken order total is much higher.

If a vendor's newer model reads raw strings correctly, you will see it in the read-back: the prepared and the raw version both pass. Then you can drop the helper for that model and keep it for the others.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume