Which live voice jobs fit Sume TTS? A four-case table
Sume TTS is submit-and-poll with no streaming. Fixed prompts and alert lines fit; a conversational turn does not. A four-case table, with costs.

Sume TTS fits live use only when you can render the line before the moment it plays: greetings, hold messages, price alerts. It does not fit a spoken reply that depends on what a caller just said, because Sume's TTS Router lists streaming TTS as a non-goal and every job is submitted, then read back from a result URL.
Two numbers that get mixed up
Microsoft's MAI-Voice-2.1 model page lists about 45 ms for the Flash model and about 550 ms for the Standard model, and positions Flash for call centers and voice assistants (read 2026-10-05). Launch coverage circulated a 150 ms end-to-end figure for Flash. The two likely measure different spans, so ask what each one includes before you plan a latency budget around either.
On Sume, a request returns a job id. With mode: sync, the API holds the request for at most wait_timeout_seconds, clamped to 30. After that it returns the queued job with polling URLs and you read GET /v1/jobs/:id/result. That is a file-delivery model, so it has no time-to-first-audio number to compare with Flash.
The four-case table
The pre-render cost below uses Sume's rate of $0.0475 per 1,000 characters and 160 characters per line. A 160-character line is 0.0076 dollars raw and bills as 1 cent each.
| Case | Text known in advance? | Fits Sume TTS? | Cost on Sume |
|---|---|---|---|
| Phone greeting and 50 fixed prompts | Yes | Yes, render once and store the files | 50 lines at 1 cent each = 50 cents |
| Stream alerts with a price in the line | Yes, a bank of variants | Yes, if you pre-render every price you will announce | 1 cent per short line |
| Scheduled announcement read at 9:00 | Yes | Yes, render the night before | 7 cents for 1,350 characters |
| Voice agent reply to what a caller just said | No | No. Use a streaming engine for this leg | Not applicable |
How to pre-render safely
Remember the 20,000-character limit per request and the 1-cent minimum. Batching 50 prompts into one job saves cents, but then you must slice by sentence timings, so keep prompts independent unless the cents matter.
- Use WAV output when you need sentence slices. Per the contract,
segmentation.emit_audioslices require wav or raw containers; mp3 returns timings without slices. - Send a stable
Idempotency-Keyper prompt so a retry never pays twice. - Store results on your side by prompt id. Each
resultcarries the audio URL and durations. - Version the bank. When a price changes, re-render only the lines that changed.
What to log
For every rendered prompt, keep the job id, the text hash, the voice id and the output container. If a prompt is later reported as wrong, you can regenerate exactly that line, which costs a cent or two, instead of re-recording a whole bank.
A price alert is the clearest example. A line that reads a price wrong is worse than silence, so compare the transcript you sent with the price in your catalog before you play it, and treat the audio as an artifact of that check.
The honest limit
If your product is a conversation, the voice engine needs to stream and the application needs to be in the loop. Use Sume for the rendered parts around it: the intro video, the hold music from Music Router, the post-call recap read, the captions burned onto the recording. Keep the live leg on an engine that was built for streaming.
How to measure before you commit
Whatever a vendor page says about milliseconds, your number is the one from your own network and your own text. Time a few dozen requests of the length you will really send, from the region where your service runs, and look at the slow end of the spread rather than the middle. A voice that answers in a fifth of a second most of the time and in two seconds one time in twenty will feel broken in a conversation.
For Sume jobs the same habit applies to the wait. A sync request waits at most 30 seconds, and a longer job returns a job id you poll. Neither is a streaming socket, so the honest question for a live feature is whether a whole utterance can be generated before the listener expects it, not whether the first sound arrives early.
Sources
Related posts
More in Use cases
- Which scripts can Sume burn into captions? Latin and Hangul, then test
Index-Translate covers 150 languages, but Sume's caption docs describe Latin and Hangul styles. Route by script and spend $0.20 on a test cue for the rest.
- Which SKUs get an AI video first? Rank by profit inside a fixed budget
Rank SKUs by 30-day gross margin and fill a fixed video budget at Sume prices: $0.625 Wan or $1.89 Seedance 2 clips plus $0.30 for captions and Timeline.
- Who said what without speaker labels: STT per track, merge by time
Sume STT has no diarization. If each speaker has an own track, transcribe each and merge words by start time into labeled turns. Python, 1 cent a minute.
- AI Shorts lost views after Oct 2, 2026? Diagnose first
Not every drop is the originality update. A diagnosis order for AI Shorts: copied clips, sameness, labels, then ordinary causes, using only what YouTube says.
Written by Sume