40 fixed phone prompts in one TTS job: no streaming needed

Fixed prompts do not need 150 ms streaming. Render them once as sentence slices with Sume TTS and play files. Worked cost for 40 prompts of 60 characters.

5 min readSume
All posts

If your phone menu says the same 40 sentences every day, you do not need a low-latency streaming voice, you need 40 audio files. Sume's TTS is an async job rather than a stream, so it fits: send all the prompts as one transcript with sentence segmentation and WAV output, and the finished job returns one gapless slice per sentence. At 60 characters each, 40 prompts is 2,400 characters, about $0.11 at the $0.0475 per 1,000 characters list rate.

When latency numbers matter and when they do not

Microsoft markets MAI-Voice-2.1-Flash for call centers, assistants and IVR, with 45 seconds of audio at 150 ms end to end and about 45 ms of model inference (read 2026-10-04). Mistral lists 70 ms to first audio for Voxtral TTS on a typical input. Those figures matter when the text is created during the call, such as a model answering a caller. They do not matter for a greeting that never changes: the file can be on disk before the phone rings.

Which prompts need a streaming voice (examples, read 2026-10-04 for vendor figures)
Prompt typeText known in advance?Fit
Greeting, hours, menu optionsYesPre-render once, play the file
Account balance with a numberPartlyPre-render the sentence frame and number words, or generate on demand
Answer from a language modelNoStreaming voice with sub-second first audio
Hold message, error tone textYesPre-render once

The request

On POST /v1/tts-1.0/generate set timestamps: { words: true } and segmentation: { mode: "sentence" }, and choose a WAV output_format. The OpenAPI file says segments come back gapless, with a 70 ms boundary lead by default, and that each segment gets a sample-exact audio_url only for wav or raw containers; with mp3 you get timings but no slices. For a phone line you may want 8000 Hz; the allowed sample rates in the schema include 8000, 16000, 24000, 44100 and 48000.

End every prompt with terminal punctuation so the sentence splitter has something to cut on, and keep the order stable, because slice index is how you map file to prompt.

Cost worked out

The public TTS price is $0.0475 per 1,000 characters, spaces and punctuation included, and the job maximum is 20,000 characters. 40 prompts of 60 characters is 2,400 characters, so the list price is 2.4 × $0.0475, about $0.114. If you submit 40 separate jobs instead, each has a minimum charge in the catalog of 1 cent, so one job is the cheaper and tidier shape. Check the live number in GET /v1/catalog before you budget.

After the job

Poll or use a webhook as described in jobs and results, download each slice, and store them by prompt id. When a prompt changes, regenerate only that sentence as a small new job; see the re-record post for the cost of a single line. If you later need to join prompts into one file, timeline audio concat joins up to 20 parts without re-synthesis.

What this does not cover: telephony itself. Sume returns files; connecting them to a phone system, handling caller input and barge-in are yours.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume