Filipino and Cantonese: fil and yue codes in MAI-Transcribe-2 vs Sume
MAI-Transcribe-2-Streaming lists fil and yue and treats unknown codes as unset. Sume takes a 2-16 character language_code hint; test a sample before relying.

If you transcribe Filipino or Cantonese, MAI-Transcribe-2-Streaming expects the three-letter codes fil and yue that appear in its 60-language list, and it treats any unrecognized value as unset, which means auto-detect. Sume STT 1.0 takes language_code as a 2 to 16 character string and falls back to auto-detect when you omit it, but Sume does not publish a per-language support table, so the answer for a given language is a short test, not a lookup.
What the two request fields accept
On the Microsoft realtime page, transcription.language is null for auto-detect and the page lists the supported codes, including fil and yue among the three-letter ones. The same page says unrecognized values are treated as unset. That is convenient and also a trap: a code like fil-PH or yue-HK that is not on the list does not error, it quietly becomes auto-detect.
In the Sume request schema, language_code is a string between 2 and 16 characters described as a BCP-47 hint, and omitting it means auto-detect. The length range allows fil, yue and regional tags; whether the provider honors a specific code is something the schema does not promise.
| Behavior | MAI-Transcribe-2-Streaming | Sume STT 1.0 |
|---|---|---|
| Field | transcription.language | language_code |
| Omitted or null | Auto-detect | Auto-detect |
| Accepted shape | Codes from the page's list, 3-letter ones included | String of 2 to 16 characters |
| Unlisted value | Treated as unset | Passed as a hint; no support table published |
| Mixed-language audio | Continuous language detection claimed | One hint per request |
A cheap check before you send
A shape check cannot tell you whether a language is supported, but it catches typos and stray region suffixes before they silently turn into auto-detect or an ignored hint. This one accepts two or three lowercase letters with an optional region and rejects the rest.
import re
TAG = re.compile(r"^[a-z]{2,3}(-[A-Z]{2})?$")
def check(tag):
if not TAG.match(tag):
raise ValueError(f"unexpected language tag: {tag!r}")
return tag
for t in ["fil", "yue", "ko", "pt-BR", "fil_PH"]:
try:
print(check(t), "ok")
except ValueError as e:
print(e)
How to settle it for your language
Run a 30 second clip of real speech through Sume with your code, then once with the hint omitted, and compare the text and the returned language_code. See the stored walkthrough on testing a language with a sample for the routine. If the two outputs agree, you can omit the hint and let detection work.
Limits: one clip proves little for dialects. Cantonese spoken with heavy English mixing, or Filipino code-switched with English, is exactly where a single hint can hurt. Test with your own audio, not a clean read.
Before you commit to a language
A short routine saves a surprise later. None of it needs more than one clip per language.
- Record or pick 30 seconds of the speech you actually expect, not a studio read.
- Transcribe once with your hint and once without, and compare the text and the returned language code.
- Write down which of the two you ship, and the clip that justified it.
- Re-run when the vendor's language list or your audio source changes; the Microsoft page is a public preview and can change.
Sources
Related posts
More in Models
- FLUX 3 Image open weights are 'coming weeks': API now or wait?
Press says FLUX 3 Image open weights arrive in the coming weeks; BFL lists commercial weights by licence. Call an API now if you need results this month.
- FLUX 3 Image combines 10 references: Sume's limits per model
BFL says FLUX 3 Image combines up to ten references. Sume lists FLUX.2, not FLUX 3: the reference ceilings per catalog model, read from the live descriptors.
- Gemini Omni Flash 1.1 on Sume: one model id, four request shapes
Sume routes gemini-omni-flash-1.1 by what you send: prompt, image_url, reference lists, or video_url. Fields, limits and the price of each in one table.
- Omni Flash took 7 image references; Omni 1.1 on Sume takes up to 10
Google's original Omni Flash allowed up to 7 image references and 3 short clips. Sume's Omni Flash 1.1 route lists up to 10 images and 3 clips of 3 s.
Written by Sume