Which live voice jobs fit Sume TTS? A four-case table

Sume TTS is submit-and-poll with no streaming. Fixed prompts and alert lines fit; a conversational turn does not. A four-case table, with costs.

5 min readSume
All posts

Sume TTS fits live use only when you can render the line before the moment it plays: greetings, hold messages, price alerts. It does not fit a spoken reply that depends on what a caller just said, because Sume's TTS Router lists streaming TTS as a non-goal and every job is submitted, then read back from a result URL.

Two numbers that get mixed up

Microsoft's MAI-Voice-2.1 model page lists about 45 ms for the Flash model and about 550 ms for the Standard model, and positions Flash for call centers and voice assistants (read 2026-10-05). Launch coverage circulated a 150 ms end-to-end figure for Flash. The two likely measure different spans, so ask what each one includes before you plan a latency budget around either.

On Sume, a request returns a job id. With mode: sync, the API holds the request for at most wait_timeout_seconds, clamped to 30. After that it returns the queued job with polling URLs and you read GET /v1/jobs/:id/result. That is a file-delivery model, so it has no time-to-first-audio number to compare with Flash.

The four-case table

The pre-render cost below uses Sume's rate of $0.0475 per 1,000 characters and 160 characters per line. A 160-character line is 0.0076 dollars raw and bills as 1 cent each.

Four live-use cases and whether job-based TTS fits (Sume TTS Router doc and Microsoft model page, read 2026-10-05)
CaseText known in advance?Fits Sume TTS?Cost on Sume
Phone greeting and 50 fixed promptsYesYes, render once and store the files50 lines at 1 cent each = 50 cents
Stream alerts with a price in the lineYes, a bank of variantsYes, if you pre-render every price you will announce1 cent per short line
Scheduled announcement read at 9:00YesYes, render the night before7 cents for 1,350 characters
Voice agent reply to what a caller just saidNoNo. Use a streaming engine for this legNot applicable

How to pre-render safely

Remember the 20,000-character limit per request and the 1-cent minimum. Batching 50 prompts into one job saves cents, but then you must slice by sentence timings, so keep prompts independent unless the cents matter.

  • Use WAV output when you need sentence slices. Per the contract, segmentation.emit_audio slices require wav or raw containers; mp3 returns timings without slices.
  • Send a stable Idempotency-Key per prompt so a retry never pays twice.
  • Store results on your side by prompt id. Each result carries the audio URL and durations.
  • Version the bank. When a price changes, re-render only the lines that changed.

What to log

For every rendered prompt, keep the job id, the text hash, the voice id and the output container. If a prompt is later reported as wrong, you can regenerate exactly that line, which costs a cent or two, instead of re-recording a whole bank.

A price alert is the clearest example. A line that reads a price wrong is worse than silence, so compare the transcript you sent with the price in your catalog before you play it, and treat the audio as an artifact of that check.

The honest limit

If your product is a conversation, the voice engine needs to stream and the application needs to be in the loop. Use Sume for the rendered parts around it: the intro video, the hold music from Music Router, the post-call recap read, the captions burned onto the recording. Keep the live leg on an engine that was built for streaming.

How to measure before you commit

Whatever a vendor page says about milliseconds, your number is the one from your own network and your own text. Time a few dozen requests of the length you will really send, from the region where your service runs, and look at the slow end of the spread rather than the middle. A voice that answers in a fifth of a second most of the time and in two seconds one time in twenty will feel broken in a conversation.

For Sume jobs the same habit applies to the wait. A sync request waits at most 30 seconds, and a longer job returns a job id you poll. Neither is a streaming socket, so the honest question for a live feature is whether a whole utterance can be generated before the listener expects it, not whether the first sound arrives early.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume