Caption a MAI-Voice-2.1 narrated video: script_text keeps spelling
Burn captions on a clip voiced by MAI-Voice-2.1 or Flash: send the video URL and script_text, and Sume times the words. $0.20 for a clip up to 60 seconds.

To caption a video whose voice came from MAI-Voice-2.1 (or Flash), host the finished video at a public HTTPS URL and call POST /v1/video-captions with that video_url and your script in script_text. Sume transcribes the audio for timing and aligns the burned-in wording to your script, so brand names appear as you typed them. The job is $0.20 for a clip up to 60 seconds.
Sume does not need to know which TTS made the voice. The route reads audio from the video, so it works with MAI-Voice, with Sume's own TTS or with a human recording.
Where the pieces come from
Microsoft's Learn page shows MAI-Voice returning an MP3 file from SSML (its example requests audio-24khz-160kbitrate-mono-mp3). You mux that audio onto your visuals with your own tool, then publish the result. Flash is documented with a 45-second audio limit on the launch post, so a 60-second video narrated by Flash is at least two synthesis calls.
Sume's alignment fails closed rather than guess: script_alignment_mismatch or script_alignment_failed come back if the script does not match what was said. That is a feature when your SSML changed the wording, for example by expanding an abbreviation.
| Approach | Fields | Timing from | Cost up to 60 s |
|---|---|---|---|
| Sume STT only | video_url | Speech recognition | $0.20 |
| Script alignment | video_url + script_text (max 8,000 chars) | Speech recognition, your wording | $0.20 |
| Your own timings | video_url + words or cues | Your data | $0.20 |
The request
Keep script_text to what is actually spoken, without SSML tags. If the clip has no speech, the job fails as caption_no_speech, and the answer is cues instead.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"video_url": "https://example.com/launch-narrated.mp4",
"script_text": "Meet Lumara Nova, the lamp that follows your day.",
"language": "en",
"style": "black-outline"}'Caveats
The video URL has to be fetchable without a login, and signed or private URLs are rejected. Captions are burned into the pixels, so keep the clean master for other languages. Microsoft marks MAI-Voice as a public preview without an SLA (read 2026-10-08); keep the audio files so you can re-voice without re-captioning from scratch.
Checks before you publish
Watch the captioned clip once with the sound off. Confirm the first and last words appear, that line breaks do not hide the product name, and that on-screen text from the video itself is not covered by the captions. If the style looks wrong for the platform, re-run with a different style or a design override instead of re-voicing the clip.
Keep the uncaptioned master and the exact script in the same folder. The captions route also accepts a source_caption_id to start from an earlier caption; check the docs page for what that reuses and what it bills.
Sources
Related posts
More in Integrations
- ENABLE_TOOL_SEARCH=false in Claude Code loads every Sume tool up front
Setting ENABLE_TOOL_SEARCH=false or a custom ANTHROPIC_BASE_URL turns off Claude Code tool search. What that does to a Sume MCP session.
- Claude Code tells Claude when an MCP server fails: Sume down vs auth
With tool search on, Claude Code reports failed MCP servers to the model. How to read that when Sume's server will not connect, 401 or 403.
- Can GLM-5.3 call Sume's hosted MCP? Function calling, not MCP
Z.ai's GLM-5.3 page lists function calling and does not mention MCP. How a GLM agent reaches Sume through an MCP client or through plain HTTP.
- How to add an MCP server to ChatGPT with developer mode
Turn on ChatGPT developer mode, create an app for the server's URL, and sign in with OAuth. The steps, with Sume's hosted MCP server as the example.
Written by Sume