MAI-Transcribe-2 languages (60) vs Sume STT language hint

Microsoft's MAI-Transcribe-2 preview covers 60 languages. Sume STT 1.0 takes one optional language_code hint and auto-detects when you omit it.

4 min readSume
All posts

Microsoft's MAI-Transcribe-2 announcement says the model, in public preview in Microsoft Foundry, expands coverage to 60 languages. Sume's speech-to-text, sume/stt-1.0, does not publish a language count. It takes one optional language_code hint, and with the field omitted it auto-detects.

Microsoft's claims are from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. This is a language-coverage reading, not an accuracy comparison.

What does Microsoft say about MAI-Transcribe-2 languages?

The post says the model "expands coverage to 60 languages" and lists "60 languages total" as one of three new capabilities, next to speaker diarization and word-level timestamps. A later use-case paragraph on meetings says "43 languages supported". The page does not reconcile the two figures, so check Microsoft's language list before you rely on a specific language.

What is Sume's language input?

The request schema describes language_code as an optional BCP-47 or provider language hint, with en and ko as examples, and says to omit it for auto-detect. The same schema allows 2 to 16 characters. There is no list of supported languages in the schema, and no multi-language field.

Language handling, Microsoft announcement vs Sume STT schema, read 2026-10-01.
ItemMAI-Transcribe-2 (preview)Sume `sume/stt-1.0`
Language coverage60 languages total (43 in one paragraph)No count published
Pick a languageNot described in the postOptional language_code hint
Leave it outNot described in the postAuto-detect
Model idMAI-Transcribe-2sume/stt-1.0; provider models stay internal

Should I set language_code or omit it?

Set it when you know the language, for example ko for a Korean interview. Omit it for mixed or unknown audio. For a language-specific walkthrough see STT language hint: Korean vs auto-detect.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-demo-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
    "duration_seconds": 120,
    "language_code": "ko"
  }'

What are the other limits on a Sume request?

duration_seconds is optional, from 1 to 600 (a maximum of 10 minutes), and it only improves the usage reservation. Omit it and Sume reserves 1 minute. For recordings longer than that, transcribe long audio files shows how.

Sources

Related posts

More in Models

All Models posts

Written by Sume