Captions for a clip that switches English and Spanish mid-sentence
MAI-Transcribe-2-Streaming advertises continuous language detection. For a recorded code-switching clip on Sume, lock the wording with script_text.

For a recorded clip where the speaker flips between English and Spanish, the dependable way to get correct captions on Sume is to supply the wording yourself. Send script_text to the video captions endpoint: Sume keeps speech-to-text word timings as the timing source and aligns your text to them. Microsoft says MAI-Transcribe-2-Streaming does automatic, continuous language detection across 60 languages (read 2026-10-04); Sume's docs describe a single language hint or automatic detection, not a per-word promise.
What each side documents
Two different tools, two different promises.
| Tool | Documented behavior | Implication for mixed speech |
|---|---|---|
| MAI-Transcribe-2-Streaming | 60 languages, automatic and continuous language detection, first partials just over 100 ms | Built for live speech that changes language |
| Sume video captions | language is a speech-to-text hint; omit for automatic detection | Treat the hint as one language per job |
Sume captions with script_text | Wording locked to your script; timings from speech | You own spelling in both languages |
| OpenAI speech-to-text | prompt, keywords and languages parameters for multilingual audio | Context can steer mixed-language audio |
The approach for a recorded clip
Write the script as spoken, with both languages in it, including accents and inverted punctuation. Then send it with the clip:
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: codeswitch-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clip.mp4",
"style": "slam",
"script_text": "Okay, this one is perfecto for the weekend. Y si lo quieres azul, dime."
}'What can go wrong
The alignment step maps your words onto the timing the speech-to-text produced. If the engine heard something very different, the job fails with script_alignment_mismatch or script_alignment_failed, and the docs suggest simplifying the script or omitting it. That failure is a feature: a mismatched caption is worse than none. Another limit is the 60-second ceiling the fixed price covers, $0.20 per accepted job for videos up to 60 seconds.
Script and speech must also fit the style. The slam style is a Latin display face; accented Latin letters are fine, but Hangul text on a Latin style is rejected, as the video captions page explains.
When you do need live detection
If the audio is a live event, a file-based caption job is the wrong shape. Sume caption jobs finish and are then fetched, and Sume webhooks deliver terminal events only, with no partial results. For live subtitles use a streaming recogniser, then send the recording to Sume afterwards for the burned-in version. Jobs and results describes how the polling side works.
Sources
Related posts
More in Developers
- Watch the Sume video catalog for new ids and changed limits in Python
Fetch GET /v1/videos/models, save a snapshot and diff new ids, removed ids and changed durations or resolutions. A Python script, testable offline.
- Chapter markers from Sume STT sentence segments
Request segmentation mode sentence on Sume STT, get time ranges per sentence, and turn the ones you pick into 0:00-style chapter lines with a 12-line formatter.
- Check Sume artifact size, width and duration before you download
Read size_bytes, width, height and duration_ms from the job result and reject an unexpected artifact before spending bandwidth. Node 18 TypeScript sample.
- Choose an image model in code from the Sume catalog's parameters
Filter GET /v1/images/models by what a request needs (references, ratio, transparency), then rank the matches by endpoint price. Python script for Sume.
Written by Sume