Make an AI voice read numbers, prices and years correctly: a test list

Test how a Sume TTS voice reads prices, years, percentages and phone numbers: eight lines in one job, sentence slices to audition, and a respell fallback.

5 min readSume
All posts

To make an AI voice read numbers, prices and years correctly, do not guess: put one of each in a short test list, generate it once, and listen. With Sume TTS you can send eight test lines as one job, get a separate audio slice per sentence, and costs under a cent. Where a line reads wrong, respell it in words, such as writing "twenty twenty-six" for a year, and re-test only that line.

The sentence slices, language field and price come from the Sume API reference and API pricing, and the polling steps from Jobs and results, read on 2026-10-03. Sume takes a plain transcript with no SSML and no documented number-format field, so respelling is the control you have. I did not generate audio for this post, so I make no claim about how any voice reads these lines.

What belongs in the test list?

Cover the shapes your scripts use. Keep one idea per line and end each line with a full stop, so sentence segmentation gives one slice per line.

Eight number shapes to test, with a respelled fallback, read 2026-10-03.
ShapeTest lineRespell if it reads wrong
Price with centsThe price is $1,299.99.one thousand two hundred ninety-nine dollars and ninety-nine cents
Small priceIt costs $0.05 per second.five cents per second
Percent and dateSave 15% before 10/31.fifteen percent before October thirty-first
Phone numberCall 555-0142 today.five five five, zero one four two
Version and yearVersion 2.5 ships in 2026.two point five ships in twenty twenty-six
TimeThe meeting is at 3:30 PM.three thirty P M
Order numberOrder #4521 left on October 3.order forty-five twenty-one
MultiplierWe grew 3x in Q4.three times in the fourth quarter

How do I send the list as one job?

Request a wav container, timestamps.words and segmentation with mode: "sentence". Segmentation needs word timings, and per-sentence audio_url slices come only for wav or raw. The script builds the body and prints the cost. After the job finishes, check that segments[] has eight entries; if a decimal or an abbreviation split a line, the count will differ and tell you where to look.

import json
RATE = 0.0475 / 1000
tests = [
    "The price is $1,299.99.",
    "It costs $0.05 per second.",
    "Save 15% before 10/31.",
    "Call 555-0142 today.",
    "Version 2.5 ships in 2026.",
    "The meeting is at 3:30 PM.",
    "Order #4521 left on October 3.",
    "We grew 3x in Q4.",
]
transcript = " ".join(tests)
body = {
    "transcript": transcript,
    "language": "en",
    "voice": {"id": "voi_0123456789abcdef0123456789abcdef"},
    "output_format": {"container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le"},
    "timestamps": {"words": True},
    "segmentation": {"mode": "sentence", "emit_audio": True},
}
print(json.dumps(body)[:120], "...")
print(len(tests), "lines,", len(transcript), "characters =", round(len(transcript) * RATE, 5), "USD")

What does it cost to re-test?

Sume bills $0.0475 per 1,000 characters, spaces and punctuation included, so the 197-character list costs about $0.009. A respelled line is usually longer, so count again before you resend. A test list is cheaper than a retake of a full voiceover with a wrong price in it.

What else should I check?

Listening finds pronunciation errors; it does not prove the words are all there. A round trip through speech-to-text can flag skipped words, as in the STT check. Send language: "en" and an English voice, so the language check cannot stop the request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume