Text to speech reading a confirmation code digit by digit

To make TTS read a code one digit at a time on Sume, write the transcript with the digits separated, pin a Sonic id, and listen before you ship it.

4 min readSume
All posts

On Sume, the safe way to get a confirmation code read one digit at a time is to write the digits separated in the transcript (for example 4 8 2 1), submit it with a pinned Sonic model id, and listen to the audio before it goes out. Sume's TTS request has no text-normalization setting, so the transcript you send is what gets spoken.

Cartesia's Sonic 3.6 notes say the model voices confirmation codes correctly without preprocessing, and its advanced guide shows an OTP read digit by digit through a normalization setting. Sume's request schema does not list that setting, so this post covers what you can do with the fields Sume does have.

What does Cartesia say about codes in Sonic 3.6?

Two statements, both read on 2026-09-30. The Sonic 3.6 page says the model follows your transcript faithfully and voices confirmation codes and heteronyms correctly without preprocessing. The advanced-capabilities guide shows a Hindi sentence containing an OTP, 4821, read the English way as four eight two one rather than as a Hindi number, by setting normalization to en-IN on Cartesia's own API.

Code reading, vendor claim against Sume request fields, read 2026-09-30.
QuestionCartesia docsSume request
Codes read correctly without preprocessingStated for Sonic 3.6Not a Sume claim; listen to the output
Digit-by-digit controlnormalization setting with a locale such as en-INNo such field in the schema
Language of the transcriptlanguage or localelanguage, 2 to 16 characters

How do I write the transcript for a code?

Sume reads the transcript literally: the MCP tool guidance says the literal transcript is the default in every thread. So put the shape you want heard in the text itself. Separate the digits with spaces or commas, and keep the sentence around them short.

Whether a given model then pauses or groups the digits as you hoped is something to hear, not assume. Sume's guidance is that the receipt proves input integrity, not pronunciation, so a successful job tells you the text arrived unchanged, not how it sounded.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: otp-demo-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Your confirmation code is 4 8 2 1.",
    "avatar_handle": "acme",
    "language": "en"
  }'

Which model id should I pin?

On the TTS Router, model is required and is a pass-through catalog id. The schema's enum lists sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. Pin sonic-3.6 explicitly if the Cartesia claim above is the reason you are choosing it; an alias such as sonic-latest can point somewhere else later. The ids are covered in Cartesia Sonic 3.6 API model ids.

Can a pronunciation dictionary help?

The TTS schema has an optional pronunciation_dict_id field described only as an optional pronunciation dictionary id. Sume's docs do not say how a dictionary treats digit strings, so test it on your own codes before relying on it. For general word fixes, see Text to speech pronunciation.

What should I check before sending codes to customers?

Generate a few real-looking codes, including ones with repeated digits and leading zeros, and listen to each. Keep the code itself out of logs if it is a live secret. For call and voicemail flows that use short fixed prompts, see IVR and voicemail prompts with text to speech.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume