Bilingual captions: continuous language detection vs a language hint
MAI-Transcribe-2-Streaming detects language continuously across 60 languages. Sume's caption `language` is a hint. What that means for bilingual clips.

If your speakers switch between two languages mid-sentence, the transcription model needs to detect language continuously, and Microsoft says MAI-Transcribe-2-Streaming does that across 60 languages (read 2026-10-03). Sume's burned-in captions work differently: the language field is a speech-to-text hint, and when you omit it the docs say detection is automatic. The docs do not describe per-phrase switching, so test a bilingual clip before you promise a client anything.
This post separates the two jobs, live transcription and burned-in captions on a finished video, and gives you a way to check both on your own audio.
What each side says
The two are not substitutes. One is a live engine you wire into a stream; the other is a render step that burns text onto an existing clip. A team that captions live and then publishes the recording would use both.
| Question | MAI-Transcribe-2-Streaming | Sume video captions |
|---|---|---|
| Language handling | Continuous language detection, 60 languages | language hint (ko, en, ...); omit for automatic detection |
| Live use | Streaming, first partials in just over 100 ms | No streaming: submit a finished public video URL |
| Price | $0.54 per hour, introductory through year end | $0.20 per job, videos up to 60 seconds |
| Word timing | Timestamps and diarisation listed | Word timings reused on a source_caption_id restyle |
| Wrong-language fix | Not stated on the pages I read | Pass words, cues or script_text to set the wording |
What a language hint does and does not do
Sume's caption docs are explicit about one thing that surprises people: language never selects the style or the font. It only tells speech-to-text what to expect. An omitted style is decided by the wording itself, slam for Latin text and black-outline for Korean. So a bilingual clip with English and Spanish gets Latin styling either way, and the hint only affects recognition.
If recognition drops one language, you have an escape hatch that does not depend on detection at all: supply the wording yourself. words (word level) or cues and segments (phrase level) skip speech-to-text and burn exactly that copy at the times you give. script_text keeps speech-to-text timings and aligns your wording to them, and fails with typed errors such as script_alignment_mismatch if it cannot.
A bilingual test you can run
Make a 40 second clip with four switches: a full sentence in each language, then two mid-sentence swaps. Submit it twice, once with no language and once with a hint for the main language, then compare the burned text with a transcript you wrote. Count the words that landed in the wrong language and the ones that are missing.
import re
def word_errors(reference, hypothesis):
ref = re.findall(r"\w+", reference.lower())
hyp = re.findall(r"\w+", hypothesis.lower())
# simple edit distance over words
prev = list(range(len(hyp) + 1))
for i, r in enumerate(ref, 1):
cur = [i]
for j, h in enumerate(hyp, 1):
cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (r != h)))
prev = cur
return prev[-1], len(ref)
ref = "welcome everyone bienvenidos a todos let's start the demo ahora"
runs = {
"no hint": "welcome everyone when venidos a todos let's start the demo ahora",
"hint en": "welcome everyone ben venidos at todos let's start the demo a hora",
}
for name, hyp in runs.items():
err, n = word_errors(ref, hyp)
print(f"{name}: {err}/{n} word errors ({err / n:.0%})")
Cost of getting it wrong
A mis-detected word in a live caption disappears in a second. A mis-burned word in a published video stays there. That asymmetry is why the finished-video path has typed correction routes and the live path has partials that the engine revises. Microsoft's own pages describe partials arriving in about 100 ms and then being finalised, so show partials in grey and burn only final text.
Budget the test honestly. A 40 second clip fits Sume's 60 second, $0.20 per job caption band, so two runs cost $0.40. For the live engine, Microsoft's $0.54 per hour introductory rate means a ten minute test is about nine cents. Neither is expensive, so test on your own audio rather than relying on a headline language count.
Decision rule
Pick on the failure you cannot afford. If a wrong word on screen is a refund, write the wording yourself and use words or cues, then pay for a human read. If speed matters more, accept the automatic path and spot check the first run. For live streams, Sume does not offer a streaming transcriber that I read in the docs, so a live caption feed needs a separate engine. The one Sume STT path I found is POST /v1/stt-1.0/transcribe, mentioned on the timeline audio page, and I did not read its price or limits.
Whatever you choose, keep the test clip. Language handling is the part vendors change most often between releases.
Sources
Related posts
More in Comparisons
- Cartesia overage $38-$65 per million credits vs Sume TTS price
Cartesia charges overage of $65, $45 or $38 per million credits by plan. Sume TTS lists $47.50 per million characters. Where each one wins, with the math.
- Claude Code mod vs MCP server vs skill vs hook: where Sume fits
Claude Code mods, MCP servers, skills and settings hooks overlap. Which one gives an agent Sume's image and video tools, and which one only guards them.
- Colossyan 20-30 minute video cap vs Sume avatar videos of 4-60 seconds
Colossyan's pricing page lists a per-video limit of 20 to 50 minutes and 40 scenes. Sume's avatar video takes 4 to 60 seconds per job.
- Comfy Agent Ask or Auto mode vs unattended Sume Format runs
Comfy Agent asks before each run or runs on its own. Sume API runs never ask. See which controls replace the approval prompt when a batch goes overnight.
Written by Sume