Pilot STT in 3 languages: language_code hint vs auto-detect, 30 clips
MAI-Transcribe-2-Streaming advertises 60 languages; here is a 30-clip pilot on Sume STT that compares a language_code hint with auto-detect for about 60 cents.

Short answer
To check whether a language hint helps, transcribe the same clips twice on Sume STT, once with language_code set and once without, then compare both with a hand-corrected transcript. Three languages with ten one-minute clips each is 30 audio minutes per pass, about 30 cents at the public $0.01 per audio minute, so both passes cost about 60 cents.
The prompt for this is the language claim in the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07): Microsoft describes MAI-Transcribe-2-Streaming as giving real-time transcripts in 60 languages with automatic, continuous language detection. A language count says a model supports a language, not how well it does on your audio.
The Sume docs describe language_code as an STT hint, with an example of en or ko, and say that auto-detect is used when it is not sent. That makes the comparison a clean one-field change.
| Item | Value |
|---|---|
| Languages | 3, chosen from your real audience |
| Clips per language | 10, about one minute each |
| Audio per pass | about 30 minutes |
| Passes | 2: hint set, hint not sent |
| Cost at $0.01 per audio minute | about $0.30 per pass, about $0.60 total |
Choose the clips
Choose clips that look like production: phone recordings, accents, names of your products. Pull a few that mix two languages, because that is where an auto-detect and a fixed hint can disagree. Send duration_seconds with each request; without it the service reserves one minute.
Score it
- Write a reference transcript by hand for each clip before you look at either result.
- Count substitutions, deletions and insertions per clip for each pass; the stored post on measuring word error rate has a script.
- Compare the passes per language, not as one total, because a hint can help one language and do nothing for another.
- Record any clip where auto-detect picked the wrong language; that is a vote for always sending the hint.
What the result means
The streaming claim and your file job are different tests, and the report's figures come from a streaming benchmark. A file pilot on your audio is the number that applies to a recorded workflow. If a hint shows no gain on any of your three languages, you save a field to maintain; if it helps one, send it for that one.
Reading the decision
Make the decision rule before you look at the scores. For example: send language_code for a language if the hinted pass has fewer errors on at least seven of the ten clips; otherwise leave it out. Fixing the rule ahead of time stops a single clip from deciding the outcome.
Ten clips per language is a small sample, and word error rates on short clips move a lot with one wrong name. Treat the result as a direction, and expand to more clips only for a language where the decision is close.
Keep audio in the same format for both passes. A hint cannot make up for a clip that is too quiet or clipped, so fix recording problems first.
- Write the decision rule first.
- Use the same files and format in both passes.
- Grow the sample only where the result is close.
Sources
Related posts
More in Comparisons
- Tavus PAL Maker, API, Enterprise vs the Sume avatar API: what matches
Tavus's site lists PAL Maker, a Developer API (CVI) and Enterprise. Which of them overlaps with Sume's avatar clips, and which does not. Pages read 2026-10-07.
- Text-to-video or image-to-video: when the prompt alone is enough
Use text-to-video when the look is open, and image-to-video when a frame, a face or a product must match. The Sume models that take each, and a decision rule.
- Can I use Suno Speech beta for an ad voiceover? What to check first
Suno Speech beta makes one track with voice and music. Before using it for ads, check price, languages and edits, then see how Sume splits voice and music.
- Zoom in on a small product in frame: AI recompose or crop and upscale
Product too small in the photo? A Pillow crop plus Sume Image Upscale ($0.20) keeps real pixels; an Ideogram 4.5 recompose edit costs $0.075 and redraws them.
Written by Sume