Muse Voice Transcribe 25+ languages and code-switching vs Sume STT
Muse Voice Transcribe supports 25+ languages and code-switching in one conversation. For mixed-language audio on Sume STT, omit language_code or set one hint.

Meta says Muse Voice Transcribe supports 25+ languages and can transcribe multilingual speech and code-switching within the same conversation. Sume's docs make no code-switching claim. For mixed-language audio on sume/stt-1.0, you either omit language_code for auto-detect or set one hint.
Meta's claims are from its developer post and guide; Sume's from the API reference and video inspect docs, read 2026-10-01.
What does Meta claim about languages?
The post lists "Multilingual support: Transcribe multilingual speech and code-switching within the same conversation", and the guide says the model works across 25+ languages. The same guide shows a languageBias option for when you already know the language.
What does Sume document?
language_code is an optional BCP-47 or provider hint, with en and ko as examples, and omitting it means auto-detect. On video inspect with transcribe: true, the same field is described as an STT hint. There is no list of languages and no field for several expected languages.
| Case | Muse Voice Transcribe | Sume `sume/stt-1.0` |
|---|---|---|
| Number of languages | 25+ | Not published |
| Mixed languages in one recording | Described as supported | No claim in the docs |
| Known single language | languageBias | language_code |
| Unknown language | Not described | Omit language_code: auto-detect |
What should I send for a bilingual clip?
Omit language_code and read the result. If the transcript is wrong in one language, split the audio at the switch and send each piece with its own hint. Sume does not document a multi-language field; compare the gpt-transcribe multiple-language post.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120
}'Where does language detection live on Sume?
Inside the STT job, with no separate switch. The tool for checking which language you got is the transcript itself; see detect language from audio for what the API exposes.
Sources
Related posts
More in Models
- Muse Voice Transcribe keyword biasing vs Sume STT: no vocab field
Muse Voice Transcribe has keyword biasing for names and domain terms. Sume STT has no vocabulary field: its request takes a language hint and word timings.
- Nano Banana multiple images at once: n range vs Gemini's count
Google says Gemini won't always return the exact image count you ask for in a prompt. On Sume, set n, and read each model's n range from the catalog.
- Nano Banana video to image: poster from a video via Sume stills
Google's Gemini API takes a video as context for a thumbnail or poster. Sume's images API takes image references only, so grab a still frame first.
- OpenAI organization verification for GPT Image 2.5 and Sume ids
OpenAI says you may need API Organization Verification before using GPT Image models. Sume lists the same two models as openai/gpt-image-2.5 and -sunburst.
Written by Sume