40 fixed phone prompts in one TTS job: no streaming needed
Fixed prompts do not need 150 ms streaming. Render them once as sentence slices with Sume TTS and play files. Worked cost for 40 prompts of 60 characters.

If your phone menu says the same 40 sentences every day, you do not need a low-latency streaming voice, you need 40 audio files. Sume's TTS is an async job rather than a stream, so it fits: send all the prompts as one transcript with sentence segmentation and WAV output, and the finished job returns one gapless slice per sentence. At 60 characters each, 40 prompts is 2,400 characters, about $0.11 at the $0.0475 per 1,000 characters list rate.
When latency numbers matter and when they do not
Microsoft markets MAI-Voice-2.1-Flash for call centers, assistants and IVR, with 45 seconds of audio at 150 ms end to end and about 45 ms of model inference (read 2026-10-04). Mistral lists 70 ms to first audio for Voxtral TTS on a typical input. Those figures matter when the text is created during the call, such as a model answering a caller. They do not matter for a greeting that never changes: the file can be on disk before the phone rings.
| Prompt type | Text known in advance? | Fit |
|---|---|---|
| Greeting, hours, menu options | Yes | Pre-render once, play the file |
| Account balance with a number | Partly | Pre-render the sentence frame and number words, or generate on demand |
| Answer from a language model | No | Streaming voice with sub-second first audio |
| Hold message, error tone text | Yes | Pre-render once |
The request
On POST /v1/tts-1.0/generate set timestamps: { words: true } and segmentation: { mode: "sentence" }, and choose a WAV output_format. The OpenAPI file says segments come back gapless, with a 70 ms boundary lead by default, and that each segment gets a sample-exact audio_url only for wav or raw containers; with mp3 you get timings but no slices. For a phone line you may want 8000 Hz; the allowed sample rates in the schema include 8000, 16000, 24000, 44100 and 48000.
End every prompt with terminal punctuation so the sentence splitter has something to cut on, and keep the order stable, because slice index is how you map file to prompt.
Cost worked out
The public TTS price is $0.0475 per 1,000 characters, spaces and punctuation included, and the job maximum is 20,000 characters. 40 prompts of 60 characters is 2,400 characters, so the list price is 2.4 × $0.0475, about $0.114. If you submit 40 separate jobs instead, each has a minimum charge in the catalog of 1 cent, so one job is the cheaper and tidier shape. Check the live number in GET /v1/catalog before you budget.
After the job
Poll or use a webhook as described in jobs and results, download each slice, and store them by prompt id. When a prompt changes, regenerate only that sentence as a small new job; see the re-record post for the cost of a single line. If you later need to join prompts into one file, timeline audio concat joins up to 20 parts without re-synthesis.
What this does not cover: telephony itself. Sume returns files; connecting them to a phone system, handling caller input and barge-in are yours.
Sources
Related posts
More in Use cases
- Re-render only the SKUs that changed: hash rows for a Sume bulk run
Retail surveys cite inventory imbalance and reactive decisions. Hash each product row, then send only changed rows to a Sume bulk queue with a revision key.
- Restaurant promo clip: keep Gemini Omni's native audio or swap it
Gemini Omni Flash 1.1 always makes synced audio. Keep it, or drop it with video trim ($0.02) and lay a Timeline soundtrack under the picture.
- SB 942 at $5,000 a day: what an API user should keep
California SB 942 reportedly became operative Aug 2, 2026 with $5,000 per violation per day. Not legal advice: the job records Sume gives you to keep.
- School closing announcement as a 30-second avatar video with captions
Turn a snow-day notice into a 30-second captioned avatar video in a 9:16 frame. Standard is the fastest Sume tier; cost by tier and a send-ready script.
Written by Sume