Spoken repair estimates in English and Spanish: 40 a month on Sume TTS

An auto shop that voices 40 estimates a month in two languages, 500 characters each, pays $1.90 on Sume TTS. MAI-Voice-2.1 lists $0.88. What the gap buys.

5 min readSume
All posts

An auto repair shop that sends a spoken version of 40 estimates a month, each about 500 characters, in both English and Spanish, produces 40,000 characters of speech. On Sume TTS at $0.0475 per 1,000 characters that is $1.90 a month. The same characters at Microsoft's launch list of $22 per million are $0.88 on MAI-Voice-2.1 and $0.60 on MAI-Voice-2.1-Flash, so the gap is $1.02 to $1.30 a month. Settle language and delivery before you argue about that.

The three bills

Microsoft announced both voices on 2026-10-01 and lists them at $22 and $15 per million characters, with 23 languages and 26 locales (Microsoft AI, read 2026-10-05). Flash makes up to 45 seconds of audio per call.

Both Microsoft voices are described by their maker as covering the same 23 languages, so Spanish is not a reason to prefer one over the other; the difference between them is price and the 45-second cap on Flash.

40 estimates x 2 languages x 500 characters = 40,000 characters (read 2026-10-05)
OptionRateMonthly cost
MAI-Voice-2.1-Flash$15 per 1M characters$0.60
MAI-Voice-2.1$22 per 1M characters$0.88
Sume TTS$0.0475 per 1,000 characters$1.90

What the Sume request needs

A Sume TTS call takes a transcript of up to 20,000 characters and a voice with an id. The id must be a voice UUID or a Voices library id, not a voice name; copy it from the voice catalog in Assets. Set language on every Spanish request, and pick a voice that speaks Spanish: the schema has a confirm_language_mismatch flag, so a language and voice mismatch is something you must confirm. Do not set it for customer-facing audio.

Spaces and punctuation count toward the character bill. The speed field takes slow, normal or fast, and a slower read suits an older customer listening on a phone. The job result records the voice, language and speed it used, per the jobs and results page, so next week's estimate can sound the same.

Dry run for one estimate

The script builds the English and Spanish bodies for one estimate and prints the character count and list cost. It reads two voice ids from your environment and sends nothing.

import json, os

RATE_PER_1000 = 0.0475
estimate = {
    "en": "Your brake pads are at 3 millimeters. We recommend replacing the front pads and resurfacing the rotors. The total is $412, ready Thursday.",
    "es": "Sus pastillas de freno tienen 3 milimetros. Recomendamos cambiar las delanteras y rectificar los discos. El total es $412, listo el jueves.",
}
voices = {"en": os.environ["VOICE_EN"], "es": os.environ["VOICE_ES"]}

for lang, text in estimate.items():
    body = {
        "transcript": text,
        "voice": {"id": voices[lang]},
        "language": lang,
        "speed": "normal",
    }
    print(json.dumps(body, ensure_ascii=False))
    print(f"{lang}: {len(text)} characters, ${len(text) / 1000 * RATE_PER_1000:.4f}")

What the cheaper price leaves out

A Sume TTS result is a hosted audio_url with a duration and optional word timings, and that URL is what Timeline audio and renders accept. If your goal is a voice note texted to the customer, you also need somewhere for the file to live. Sume's file is already hosted; a file from another service has to be stored by you. Count that work before you call $1.30 a month a saving.

Use Flash only after you check the 45-second cap against your script. A 500-character estimate is short, but a long repair list may not fit one call, and you would split it. On Sume, split a long script into several requests under the 20,000-character cap and join them with Timeline audio.

Keep the numbers out of the model

A spoken estimate carries the part names, the total and the pickup day, and those come from your shop software, not from a model. Build the transcript from a template with the fields filled in by code, then send that string to TTS. The model only reads what you give it, and the character count is known before you pay: the dry run above shows 135 to 140 characters for a short estimate, and a long repair list will be several times that.

Write dollar amounts the way people say them in each language if the voice stumbles on a symbol, and test one sample first. Sume has no SSML field, so spelling a number out in the transcript is the control you have.

A pilot week

Voice ten real estimates in both languages and have a Spanish speaker listen before any goes to a customer. Check part names, prices and the pickup day, because a number read wrongly is worse than none. Keep the written estimate beside the audio, since the voice note is a convenience, not the record.

Also decide what happens when a customer replies by voice. Sume speech-to-text is $0.01 per audio minute, so a short reply can be transcribed for about a cent and read by the service writer, which keeps the whole exchange in text on the work order.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume