Does Sume STT detect the language? Omit language_code

Sume STT detects a clip's language when you omit language_code. MAI-Transcribe-2-Streaming claims continuous detection. They are not the same promise.

4 min readSume
All posts

Yes. In Sume STT 1.0, language_code is optional and the docs say to omit it for auto-detect. You can also pass a hint such as en or ko, from 2 to 16 characters. The job returns text and words[], and language fields when the provider supplies them. That is detection of the clip's language, not a promise to follow a speaker who switches languages mid-sentence.

Detection versus continuous detection

Microsoft's post says MAI-Transcribe-2-Streaming covers 60 languages with automatic, continuous language detection. The word continuous is the point: a stream can notice a change while it runs. Sume's schema describes one optional hint for the file, and says nothing about switching inside it. Test a code-switching clip on your own audio before you rely on either.

When to send a hint

Use a hint when you know the language. It removes one source of error, and it is the only way to be sure a short, noisy clip is read as the right language. Omit it for a batch of mixed files, then read the language field of each result and rerun the ones that look wrong with a hint.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-lang-hint-001" \
  -d '{"audio_url": "'"$AUDIO_URL"'", "language_code": "ko", "duration_seconds": 90}'

The two behaviours, read 2026-10-06:

Language detection, read 2026-10-06
QuestionMAI-Transcribe-2-Streaming (vendor)Sume STT 1.0 (docs)
Languages60One optional hint, 2 to 16 characters
DetectionAutomatic, continuousAutomatic when language_code is omitted
Mid-clip switchClaimed by the vendorNot documented; test it
ModeStreamingBatch job, up to 600 s

A practical test

Make three short clips: one in each of your two languages and one that mixes them. Send each without a hint, then with a hint, and compare the text. Keep the clips and the results as your own evidence.

The mixed clip is the one that matters if your audience switches languages. If it fails, split the audio at the switch and send each part with its own hint, then join the transcripts using the start offsets of the pieces.

Keep the results of the test with the clips, and note the date. Models change, and a result from one month is not a promise for the next.

For captions on a mixed-language video, a separate decision applies. Burn captions from script_text you have checked, so the words are right whatever the recognizer thought the language was.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume