MAI-Transcribe-2 languages (60) vs Sume STT language hint
Microsoft's MAI-Transcribe-2 preview covers 60 languages. Sume STT 1.0 takes one optional language_code hint and auto-detects when you omit it.

Microsoft's MAI-Transcribe-2 announcement says the model, in public preview in Microsoft Foundry, expands coverage to 60 languages. Sume's speech-to-text, sume/stt-1.0, does not publish a language count. It takes one optional language_code hint, and with the field omitted it auto-detects.
Microsoft's claims are from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. This is a language-coverage reading, not an accuracy comparison.
What does Microsoft say about MAI-Transcribe-2 languages?
The post says the model "expands coverage to 60 languages" and lists "60 languages total" as one of three new capabilities, next to speaker diarization and word-level timestamps. A later use-case paragraph on meetings says "43 languages supported". The page does not reconcile the two figures, so check Microsoft's language list before you rely on a specific language.
What is Sume's language input?
The request schema describes language_code as an optional BCP-47 or provider language hint, with en and ko as examples, and says to omit it for auto-detect. The same schema allows 2 to 16 characters. There is no list of supported languages in the schema, and no multi-language field.
| Item | MAI-Transcribe-2 (preview) | Sume `sume/stt-1.0` |
|---|---|---|
| Language coverage | 60 languages total (43 in one paragraph) | No count published |
| Pick a language | Not described in the post | Optional language_code hint |
| Leave it out | Not described in the post | Auto-detect |
| Model id | MAI-Transcribe-2 | sume/stt-1.0; provider models stay internal |
Should I set language_code or omit it?
Set it when you know the language, for example ko for a Korean interview. Omit it for mixed or unknown audio. For a language-specific walkthrough see STT language hint: Korean vs auto-detect.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"language_code": "ko"
}'What are the other limits on a Sume request?
duration_seconds is optional, from 1 to 600 (a maximum of 10 minutes), and it only improves the usage reservation. Omit it and Sume reserves 1 minute. For recordings longer than that, transcribe long audio files shows how.
Sources
Related posts
More in Models
- MAI-Transcribe-2 speaker diarization vs Sume STT: no speaker field
MAI-Transcribe-2 attributes each segment to a distinct speaker. Sume STT has no diarize field: provider knobs are fixed server-side and you get words[].
- Midjourney edit model image references: 4 vs Sume's ranges
Midjourney's V8.2 edit model takes up to 4 image references. On Sume the ceiling is per model: GPT Image 2.5 takes up to 16. Read the descriptor first.
- MiniMax-H3 Fun ControlNet Union 2.0: 8 conditions vs Sume references
MiniMax-H3-Fun-Controlnet-Union-2.0 adds Scribble, Layout and Gray to five older conditions. Sume's minimax-h3 ids take image, video and audio references.
- MiniMax H3 license for an EU company: contact MiniMax first
MiniMax H3 license Section II invites people in the EU, UK, Korea and USA to contact MiniMax about a license. What it promises and omits.
Written by Sume