Caption a MAI-Voice-2.1 narrated video: script_text keeps spelling

Burn captions on a clip voiced by MAI-Voice-2.1 or Flash: send the video URL and script_text, and Sume times the words. $0.20 for a clip up to 60 seconds.

5 min readSume
All posts

To caption a video whose voice came from MAI-Voice-2.1 (or Flash), host the finished video at a public HTTPS URL and call POST /v1/video-captions with that video_url and your script in script_text. Sume transcribes the audio for timing and aligns the burned-in wording to your script, so brand names appear as you typed them. The job is $0.20 for a clip up to 60 seconds.

Sume does not need to know which TTS made the voice. The route reads audio from the video, so it works with MAI-Voice, with Sume's own TTS or with a human recording.

Where the pieces come from

Microsoft's Learn page shows MAI-Voice returning an MP3 file from SSML (its example requests audio-24khz-160kbitrate-mono-mp3). You mux that audio onto your visuals with your own tool, then publish the result. Flash is documented with a 45-second audio limit on the launch post, so a 60-second video narrated by Flash is at least two synthesis calls.

Sume's alignment fails closed rather than guess: script_alignment_mismatch or script_alignment_failed come back if the script does not match what was said. That is a feature when your SSML changed the wording, for example by expanding an abbreviation.

Three ways to caption a MAI-Voice clip on Sume (read 2026-10-08)
ApproachFieldsTiming fromCost up to 60 s
Sume STT onlyvideo_urlSpeech recognition$0.20
Script alignmentvideo_url + script_text (max 8,000 chars)Speech recognition, your wording$0.20
Your own timingsvideo_url + words or cuesYour data$0.20

The request

Keep script_text to what is actually spoken, without SSML tags. If the clip has no speech, the job fails as caption_no_speech, and the answer is cues instead.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video_url": "https://example.com/launch-narrated.mp4",
       "script_text": "Meet Lumara Nova, the lamp that follows your day.",
       "language": "en",
       "style": "black-outline"}'

Caveats

The video URL has to be fetchable without a login, and signed or private URLs are rejected. Captions are burned into the pixels, so keep the clean master for other languages. Microsoft marks MAI-Voice as a public preview without an SLA (read 2026-10-08); keep the audio files so you can re-voice without re-captioning from scratch.

Checks before you publish

Watch the captioned clip once with the sound off. Confirm the first and last words appear, that line breaks do not hide the product name, and that on-screen text from the video itself is not covered by the captions. If the style looks wrong for the platform, re-run with a different style or a design override instead of re-voicing the clip.

Keep the uncaptioned master and the exact script in the same folder. The captions route also accepts a source_caption_id to start from an earlier caption; check the docs page for what that reuses and what it bills.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume