Pilot STT in 3 languages: language_code hint vs auto-detect, 30 clips

MAI-Transcribe-2-Streaming advertises 60 languages; here is a 30-clip pilot on Sume STT that compares a language_code hint with auto-detect for about 60 cents.

5 min readSume
All posts

Short answer

To check whether a language hint helps, transcribe the same clips twice on Sume STT, once with language_code set and once without, then compare both with a hand-corrected transcript. Three languages with ten one-minute clips each is 30 audio minutes per pass, about 30 cents at the public $0.01 per audio minute, so both passes cost about 60 cents.

The prompt for this is the language claim in the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07): Microsoft describes MAI-Transcribe-2-Streaming as giving real-time transcripts in 60 languages with automatic, continuous language detection. A language count says a model supports a language, not how well it does on your audio.

The Sume docs describe language_code as an STT hint, with an example of en or ko, and say that auto-detect is used when it is not sent. That makes the comparison a clean one-field change.

Pilot design and cost at the public STT rate (read 2026-10-07)
ItemValue
Languages3, chosen from your real audience
Clips per language10, about one minute each
Audio per passabout 30 minutes
Passes2: hint set, hint not sent
Cost at $0.01 per audio minuteabout $0.30 per pass, about $0.60 total

Choose the clips

Choose clips that look like production: phone recordings, accents, names of your products. Pull a few that mix two languages, because that is where an auto-detect and a fixed hint can disagree. Send duration_seconds with each request; without it the service reserves one minute.

Score it

  • Write a reference transcript by hand for each clip before you look at either result.
  • Count substitutions, deletions and insertions per clip for each pass; the stored post on measuring word error rate has a script.
  • Compare the passes per language, not as one total, because a hint can help one language and do nothing for another.
  • Record any clip where auto-detect picked the wrong language; that is a vote for always sending the hint.

What the result means

The streaming claim and your file job are different tests, and the report's figures come from a streaming benchmark. A file pilot on your audio is the number that applies to a recorded workflow. If a hint shows no gain on any of your three languages, you save a field to maintain; if it helps one, send it for that one.

Reading the decision

Make the decision rule before you look at the scores. For example: send language_code for a language if the hinted pass has fewer errors on at least seven of the ten clips; otherwise leave it out. Fixing the rule ahead of time stops a single clip from deciding the outcome.

Ten clips per language is a small sample, and word error rates on short clips move a lot with one wrong name. Treat the result as a direction, and expand to more clips only for a language where the decision is close.

Keep audio in the same format for both passes. A hint cannot make up for a clip that is too quiet or clipped, so fix recording problems first.

  • Write the decision rule first.
  • Use the same files and format in both passes.
  • Grow the sample only where the result is close.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume