Captions for a clip that switches English and Spanish mid-sentence

MAI-Transcribe-2-Streaming advertises continuous language detection. For a recorded code-switching clip on Sume, lock the wording with script_text.

6 min readSume
All posts

For a recorded clip where the speaker flips between English and Spanish, the dependable way to get correct captions on Sume is to supply the wording yourself. Send script_text to the video captions endpoint: Sume keeps speech-to-text word timings as the timing source and aligns your text to them. Microsoft says MAI-Transcribe-2-Streaming does automatic, continuous language detection across 60 languages (read 2026-10-04); Sume's docs describe a single language hint or automatic detection, not a per-word promise.

What each side documents

Two different tools, two different promises.

Language handling on each side (read 2026-10-04)
ToolDocumented behaviorImplication for mixed speech
MAI-Transcribe-2-Streaming60 languages, automatic and continuous language detection, first partials just over 100 msBuilt for live speech that changes language
Sume video captionslanguage is a speech-to-text hint; omit for automatic detectionTreat the hint as one language per job
Sume captions with script_textWording locked to your script; timings from speechYou own spelling in both languages
OpenAI speech-to-textprompt, keywords and languages parameters for multilingual audioContext can steer mixed-language audio

The approach for a recorded clip

Write the script as spoken, with both languages in it, including accents and inverted punctuation. Then send it with the clip:

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: codeswitch-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clip.mp4",
    "style": "slam",
    "script_text": "Okay, this one is perfecto for the weekend. Y si lo quieres azul, dime."
  }'

What can go wrong

The alignment step maps your words onto the timing the speech-to-text produced. If the engine heard something very different, the job fails with script_alignment_mismatch or script_alignment_failed, and the docs suggest simplifying the script or omitting it. That failure is a feature: a mismatched caption is worse than none. Another limit is the 60-second ceiling the fixed price covers, $0.20 per accepted job for videos up to 60 seconds.

Script and speech must also fit the style. The slam style is a Latin display face; accented Latin letters are fine, but Hangul text on a Latin style is rejected, as the video captions page explains.

When you do need live detection

If the audio is a live event, a file-based caption job is the wrong shape. Sume caption jobs finish and are then fetched, and Sume webhooks deliver terminal events only, with no partial results. For live subtitles use a streaming recogniser, then send the recording to Sume afterwards for the burned-in version. Jobs and results describes how the polling side works.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume