Does Sume STT detect the language? Omit language_code
Sume STT detects a clip's language when you omit language_code. MAI-Transcribe-2-Streaming claims continuous detection. They are not the same promise.

Yes. In Sume STT 1.0, language_code is optional and the docs say to omit it for auto-detect. You can also pass a hint such as en or ko, from 2 to 16 characters. The job returns text and words[], and language fields when the provider supplies them. That is detection of the clip's language, not a promise to follow a speaker who switches languages mid-sentence.
Detection versus continuous detection
Microsoft's post says MAI-Transcribe-2-Streaming covers 60 languages with automatic, continuous language detection. The word continuous is the point: a stream can notice a change while it runs. Sume's schema describes one optional hint for the file, and says nothing about switching inside it. Test a code-switching clip on your own audio before you rely on either.
When to send a hint
Use a hint when you know the language. It removes one source of error, and it is the only way to be sure a short, noisy clip is read as the right language. Omit it for a batch of mixed files, then read the language field of each result and rerun the ones that look wrong with a hint.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-lang-hint-001" \
-d '{"audio_url": "'"$AUDIO_URL"'", "language_code": "ko", "duration_seconds": 90}'The two behaviours, read 2026-10-06:
| Question | MAI-Transcribe-2-Streaming (vendor) | Sume STT 1.0 (docs) |
|---|---|---|
| Languages | 60 | One optional hint, 2 to 16 characters |
| Detection | Automatic, continuous | Automatic when language_code is omitted |
| Mid-clip switch | Claimed by the vendor | Not documented; test it |
| Mode | Streaming | Batch job, up to 600 s |
A practical test
Make three short clips: one in each of your two languages and one that mixes them. Send each without a hint, then with a hint, and compare the text. Keep the clips and the results as your own evidence.
The mixed clip is the one that matters if your audience switches languages. If it fails, split the audio at the switch and send each part with its own hint, then join the transcripts using the start offsets of the pieces.
Keep the results of the test with the clips, and note the date. Models change, and a result from one month is not a promise for the next.
For captions on a mixed-language video, a separate decision applies. Burn captions from script_text you have checked, so the words are right whatever the recognizer thought the language was.
Sources
Related posts
More in Comparisons
- ffmpeg script or Sume Timeline for a weekly vertical series
Compare a self-hosted ffmpeg script with Sume Timeline 1.0 for a weekly 9:16 series: what you maintain, what the API refuses, price per output minute.
- Groq Orpheus TTS: $22 per million characters and a 200-character cap
Groq lists Orpheus V1 English at $22.00 per million characters with input kept under 200 characters. Request-count and cost math against Sume TTS.
- HeyGen Pro 4K export at $49 vs Sume video upscale per clip
HeyGen's page puts 4K export on Pro at $49 a month and 1080p on Creator at $29. Sume Video Upscale charges $0.009 per input second. Here is the arithmetic.
- How much more does 4K AI video cost than 720p? Multiplier table
Veo 3.1 Fast 4K costs 3 times its 720p rate on Google's page, Omni on Sume is also 3 times, Veo Standard 1.5 times and Kling Video v3 Pro 1 times.
Written by Sume