MAI streaming regions vs Sume: no region field in the STT request
MAI-Transcribe-2-Streaming lists swedencentral, centralus, southindia (eastus2 soon). Sume's STT request has no region field, so ask support about residency.

If a region matters to you, MAI-Transcribe-2-Streaming gives you a choice: its realtime page lists swedencentral, centralus and southindia as available, with eastus2 coming soon, and the endpoint is built from your resource name. Sume STT 1.0 gives no such choice in the request: its body takes audio_url, language_code, duration_seconds, segmentation and metadata, and nothing that selects where the work runs. For a residency requirement, that is a question for Sume support, not a field you can set.
What the Microsoft pages list
The streaming page names the three available regions and the one coming soon. The voices page for MAI-Voice-2.1 lists 14 serving regions: francecentral, eastasia, southeastasia, eastus, canadacentral, eastus2, westus, westeurope, northeurope, westus2, westus3, centralindia, swedencentral and japaneast. The two lists differ, so a combined speech-in, speech-out design may not fit one region.
Both pages mark the models as public preview with no SLA, so the lists can change.
| Service | Region choice | Where stated |
|---|---|---|
| MAI-Transcribe-2-Streaming | swedencentral, centralus, southindia; eastus2 coming soon | Realtime page |
| MAI-Voice-2.1 | 14 serving regions | Voices page |
| Sume STT 1.0 | No region field in the request | Request schema |
| Sume TTS 1.0 | No region field in the request | Request schema |
What to do with that
Do not infer a Sume region from the absence of a field. The schema simply does not expose one, and nothing in the docs states a processing location, so a claim either way would be a guess. If you need an answer for a contract, ask in writing and keep the reply with your records.
Limits: this post is about control, not about legal sufficiency. A region list tells you where a vendor can run a model, not what happens to audio after transcription. Read the vendor's data terms for that; the jobs and results docs describe how Sume jobs are polled.
Questions to answer before choosing
Residency is a legal question first, so write down the requirement before comparing endpoints.
- Where must the audio be processed, and does the contract name a region or only a country?
- Is audio retained after transcription, and for how long, according to the vendor's own terms?
- Does a failover region fall inside the same boundary?
- Who will sign off, and what document do they need from the vendor?
Latency and region
Region also affects delay. A WebSocket from a user in Mumbai to a resource in swedencentral crosses continents, which adds round-trip time to a service whose headline claim is first partials in just over 100 ms. That claim does not say where it was measured, so test from the place your users are. Sume STT is a file job, so network distance affects upload and polling rather than a live conversation.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 languages vs Sume voice tags: 14 shared, 9 and 2 apart
Microsoft lists 23 TTS languages; Sume's voice library tags 16. Fourteen overlap, nine are MAI-only, two (Japanese, Tagalog) are Sume-only. The full split.
- Flash TTS '55% faster inference' vs the end-to-end time of a TTS job
Microsoft quotes 55% faster inference and 150 ms end to end for MAI-Voice-2.1-Flash. Sume TTS is an async job; here is what you can actually time.
- MAI-Voice-2.1 at $22 per million characters vs Eleven v4
MAI-Voice-2.1 is $22 per 1M characters and Flash $15; Eleven v4 is $22 per 1M in a promo through Oct 12. Sume's Sonic route is $47.50 per 1M characters.
- MAI-Voice styles via express-as vs Sume's emotion field
MAI voices set styles with SSML mstts:express-as; some only have neutral. Sume has no SSML; generation_config.emotion is a free string up to 64 characters.
Written by Sume