Live voice agent or async TTS job? When Sume TTS is the wrong tool

Microsoft's MAI-Voice and transcription models list live paths like Azure Voice Live. Sume TTS and STT are async file jobs: fine for narration, not calls.

4 min readSume
All posts

The rule

If a person is waiting on the other end of a live conversation, use a streaming voice service. If you are producing a file to publish, use a job. Sume's TTS and STT are jobs: you submit, you get a job id, and you fetch a finished audio file or a transcript. They are the right fit for narration, captions and dubbing, and the wrong fit for a phone call.

Microsoft's announcement (read 2026-10-04) describes MAI-Transcribe-2-Streaming over WebSocket with first partial results in just over 100 ms, and lists MAI-Voice-2.1-Flash at 150 ms end to end. Those numbers are for live use; Sume does not publish a latency promise for jobs.

Which shape for which task

Match the shape of the service to the shape of the task before you compare prices.

Live voice versus job-based audio (read 2026-10-04)
TaskNeedsFits
Phone agent that answers callersPartial transcripts and speech in about a secondA streaming service such as the Microsoft models above
Narration for a 90-second videoA finished file you can reviewSume TTS job
Captions for an uploaded clipWord times for a file you already haveSume STT job
Live subtitles on a streamPartial results while audio arrivesA streaming transcription service

What a job gives you instead

A Sume job has a stable id, an Idempotency-Key so a retry cannot bill twice, a status you can poll, and a result URL that stays valid. In webhook mode it also sends one signed terminal event. Those properties matter for production pipelines: you can audit exactly what was generated and when. The jobs guide lists the modes, and sync waits at most 30 seconds, which is a bounded wait for a short file, not a stream.

Do not try to build a live agent by chaining jobs. The delay of submit, poll and fetch adds up on each turn, and a caller will hear it.

A hybrid that works

Many products need both. Use a streaming service for the live call, then send the recording to a batch job afterwards for a clean transcript with word times. Or generate the fixed lines of a voice agent, greetings, hold messages and confirmations, as files in advance with TTS jobs, and let the live service handle only the open parts.

  • Pre-generate anything you say more than once.
  • Use async jobs for post-call transcripts.
  • Test the live part for delay on a real network, not on your laptop.
  • Check Microsoft's current pages; the models are in public preview.

Questions to ask before you pick

How long can the listener wait? If the answer is under two seconds, you need streaming. If it is minutes or hours, a job is simpler and easier to audit. Does the output need review? Narration, captions and dubs usually do, and a job leaves you a file to review before anything is published. Do you need to interrupt the speaker? Barge-in only works on a live connection.

Then check availability. Microsoft's announcement lists the new models on Foundry and other platforms and says LiveKit support is coming soon, so the live path is still being assembled. Treat anything marked preview as likely to change and keep your file-based pipeline independent of it.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume