Live voice agent or async TTS job? When Sume TTS is the wrong tool
Microsoft's MAI-Voice and transcription models list live paths like Azure Voice Live. Sume TTS and STT are async file jobs: fine for narration, not calls.

The rule
If a person is waiting on the other end of a live conversation, use a streaming voice service. If you are producing a file to publish, use a job. Sume's TTS and STT are jobs: you submit, you get a job id, and you fetch a finished audio file or a transcript. They are the right fit for narration, captions and dubbing, and the wrong fit for a phone call.
Microsoft's announcement (read 2026-10-04) describes MAI-Transcribe-2-Streaming over WebSocket with first partial results in just over 100 ms, and lists MAI-Voice-2.1-Flash at 150 ms end to end. Those numbers are for live use; Sume does not publish a latency promise for jobs.
Which shape for which task
Match the shape of the service to the shape of the task before you compare prices.
| Task | Needs | Fits |
|---|---|---|
| Phone agent that answers callers | Partial transcripts and speech in about a second | A streaming service such as the Microsoft models above |
| Narration for a 90-second video | A finished file you can review | Sume TTS job |
| Captions for an uploaded clip | Word times for a file you already have | Sume STT job |
| Live subtitles on a stream | Partial results while audio arrives | A streaming transcription service |
What a job gives you instead
A Sume job has a stable id, an Idempotency-Key so a retry cannot bill twice, a status you can poll, and a result URL that stays valid. In webhook mode it also sends one signed terminal event. Those properties matter for production pipelines: you can audit exactly what was generated and when. The jobs guide lists the modes, and sync waits at most 30 seconds, which is a bounded wait for a short file, not a stream.
Do not try to build a live agent by chaining jobs. The delay of submit, poll and fetch adds up on each turn, and a caller will hear it.
A hybrid that works
Many products need both. Use a streaming service for the live call, then send the recording to a batch job afterwards for a clean transcript with word times. Or generate the fixed lines of a voice agent, greetings, hold messages and confirmations, as files in advance with TTS jobs, and let the live service handle only the open parts.
- Pre-generate anything you say more than once.
- Use async jobs for post-call transcripts.
- Test the live part for delay on a real network, not on your laptop.
- Check Microsoft's current pages; the models are in public preview.
Questions to ask before you pick
How long can the listener wait? If the answer is under two seconds, you need streaming. If it is minutes or hours, a job is simpler and easier to audit. Does the output need review? Narration, captions and dubs usually do, and a job leaves you a file to review before anything is published. Do you need to interrupt the speaker? Barge-in only works on a live connection.
Then check availability. Microsoft's announcement lists the new models on Foundry and other platforms and says LiveKit support is coming soon, so the live path is still being assembled. Treat anything marked preview as likely to change and keep your file-based pipeline independent of it.
Sources
Related posts
More in Comparisons
- xAI batch mixes chat, image and video in one file; Sume uses queues
xAI's Batch API now takes chat, image and video requests in one JSONL file. A Sume bulk queue targets exactly one Format. How to split a holiday job.
- xAI batch video URLs expire in 1 hour: size the download worker
xAI's release notes say image and video URLs in batch results expire after 1 hour. How many parallel downloads you need, and how Sume's durable URLs differ.
- xAI file URLs auto-expire; Sume fetches image_url once at create
xAI's Files API can serve public URLs that auto-expire. Sume fetches an attachment image_url at run creation and copies it. What that means for hosting.
- Omni is free in YouTube Shorts: when you need the API instead
Google rolls Omni Flash out free in YouTube Shorts and Create and to Gemini subscribers; the API bills per second. Where the free tools stop and an API starts.
Written by Sume