MAI-Voice-2.1-Flash lists 45 ms; Sume TTS is an async job
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms and MAI-Voice-2.1 at about 550 ms. Sume TTS 1.0 is non-streaming: a job with a poll URL or webhook.

Microsoft lists MAI-Voice-2.1-Flash at about 45 ms latency and MAI-Voice-2.1 at about 550 ms. Those figures suit a voice agent that speaks as it thinks. Sume TTS 1.0 does not stream: the OpenAPI text says "Phase 1 is async job + poll/webhook (non-streaming)". It fits narration and video, not a live call.
Microsoft's numbers are from its MAI-Voice page, read 2026-10-01; the page does not say what the latency measures, so treat it as a vendor-stated figure. Sume's are from the API reference.
What does a Sume job do instead of streaming?
It returns a job. With mode: "async" the first response carries a status_url, result_url, events_url and cancel_url. You poll status, or pass a webhook_url and receive signed terminal events (job.completed, job.failed, job.canceled); there are no partial or progress callbacks.
Can I make it feel instant with sync mode?
Partly. mode: "sync" waits up to wait_timeout_seconds, at most 30. If the job is not done by then, the response is still successful and returns the queued or processing state, so keep polling status_url instead of resubmitting. This bounds the HTTP wait, not the job.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: line-001" \
-d '{
"transcript": "Your order has shipped.",
"avatar_handle": "@narrator",
"mode": "sync",
"wait_timeout_seconds": 20
}'Which use cases fit each?
| Use case | Better fit |
|---|---|
| Live call or agent turn | A streaming engine like Flash |
| Narration for a video | Sume async job |
| Batch of voiceover lines | Sume async jobs with Idempotency-Key |
| Per-sentence timings for captions | Sume timestamps.words and sentence segments |
What if I need both?
Use a streaming vendor for live turns and Sume for rendered output. Do not wait on Sume inside a live conversation; its own docs say clients needing a longer wait should submit async and poll from their side.
Sources
Related posts
More in Developers
- Make AI Agent fallback connection retries once: key Sume calls
Make now retries an AI Agent run once on a fallback connection. Keep the Sume Idempotency-Key out of the model's hands so a retried run cannot bill twice.
- Make a voice louder than the music: gain_db and duck_db ranges
In Timeline 1.0, raise the voice with audio.gain_db (-60 to 12) and lower the music bed with soundtrack.duck_db (0 to 20). Ranges and refusal codes.
- Mastra MCP 2.0 is 2026-07-28 only: which line for Sume?
@mastra/mcp 2.0.0 drops the legacy handshake and speaks MCP 2026-07-28 only. The Sume server defaults to 2025-11-25, so pin the version you test.
- Mastra MCPClient and JSON Schema 2019-09: Sume tool schemas
Mastra MCP 2.1.1 accepts tool inputs in JSON Schema 2019-09. Sume builds its tool inputs as plain object schemas, and tools_schema returns each contract.
Written by Sume