MAI-Transcribe-2-Streaming runs in 4 regions, MAI-Voice in 14
Microsoft lists four regions for MAI-Transcribe-2-Streaming and 14 for MAI-Voice-2.1; three overlap. Sume's STT and TTS requests have no region field.

A voice agent that pairs MAI-Transcribe-2-Streaming with MAI-Voice-2.1 can sit in three regions where both models are served: East US 2, Sweden Central and Southeast Asia. Microsoft's Learn pages list four regions for the streaming transcriber and 14 for the voice models, and only those three appear on both lists (read 2026-10-08).
This matters because a speech-in step and a speech-out step in different regions add a network hop to every turn. Pick the shared region first, then place your language model call next to it.
The two region lists
MAI-Transcribe-2-Streaming: Sweden Central, Central US, East US 2 and Southeast Asia. Both pages say the models can be accessed globally and that Azure routes requests to these serving regions, so a request may not run in the region of your resource. I did not find a page that says where a given request is served, so test from the region you deploy in.
MAI-Voice-2.1 and MAI-Voice-2.1-Flash: France Central, East Asia, Southeast Asia, East US, Canada Central, East US 2, West US, West Europe, North Europe, West US 2, West US 3, Central India, Sweden Central and Japan East.
| Region | MAI-Transcribe-2-Streaming | MAI-Voice-2.1 and Flash |
|---|---|---|
| East US 2 | Yes | Yes |
| Sweden Central | Yes | Yes |
| Southeast Asia | Yes | Yes |
| Central US | Yes | No |
| East US, West US, West Europe, Japan East and 6 others | No | Yes |
| Total regions listed | 4 | 14 |
What this means for a build
If your callers are in Europe, Sweden Central is the only European region shared by the pair. West Europe, North Europe and France Central serve the voice but not the streaming transcriber. If your callers are in Japan or India, the voice has a nearby region (Japan East, Central India) and the transcriber does not, so the audio-in leg crosses the sea.
Sume has no region choice
Sume's STT 1.0 and TTS 1.0 request schemas have no region field. You call api.sume.com and the job runs wherever Sume dispatches it. The routes are asynchronous: mode can be async, sync (a bounded wait of up to 30 seconds), subscribe or webhook, so latency is measured per job, not per audio chunk.
That suits recorded work, such as transcribing a day of calls overnight at about one cent per audio minute, or pre-rendering prompts. It does not suit a live agent. If region pinning is a requirement for your data policy, a Sume job is not the answer today; say so in your vendor review rather than discover it later.
Questions to put to a vendor
Ask where a request is processed, not only where the resource lives. Ask whether prompts, audio or transcripts are retained and for how long, and whether the answer is different for a preview model. Ask what happens to a request if the nearest serving region is down: a router that silently crosses an ocean adds latency your agent will feel as dead air.
Write the answers into your architecture notes with the date. Both Microsoft pages say the models are accessible globally and routed to serving regions, which is convenient but means region is a property of the service, not a setting you control.
Sources
Related posts
More in Models
- MAI-Voice-2.1 languages: 23 listed, no Japanese or Arabic voice
Microsoft's MAI-Voice-2.1 voice table covers 23 languages and 28 locale codes, with no Japanese or Arabic row. How Sume's TTS language field handles both.
- MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.
- Nano Banana 2.1 vs Pro from 512 to 4K: price per image on Sume
Nano Banana 2.1 runs $0.075 at 512 to $0.20 at 4K; Nano Banana Pro is $0.1875 up to 2K and $0.375 at 4K. Full tier table and a per-1,000 view.
- Same Omni video twice: no seed, but a replay returns it
Omni Flash lists no temperature or seed, and Sume rejects seed on every video model. To repeat a clip, replay an idempotency key or use a video_url edit.
Written by Sume