Convert MAI-Transcribe-2 word offsets to Sume video caption words
Turn word timings in milliseconds into the seconds-based words array Sume video-captions accepts. A short Python converter plus the 60 s and 1,200-word limits.

If you transcribe with MAI-Transcribe-2 and want the words burned onto a short video, convert each word's millisecond offset and duration into start and end in seconds and send the list as words to POST /v1/video-captions. When words is present Sume skips its own speech-to-text and burns exactly those words at exactly those times. The job costs $0.20 for a clip up to 60 seconds.
This post assumes the Azure Speech batch response shape, in which each phrase has a words list and every word has text, offsetMilliseconds and durationMilliseconds. Check your own response before relying on those names; Microsoft's launch post for MAI-Transcribe-2 (read 2026-10-08) does not document the field layout.
The limits you convert into
Sume's words entries take text (1 to 200 characters), and start and end as numbers from 0 to 60. The array holds at most 1,200 entries and cannot be combined with script_text, cues or segments. A clip over 60 seconds is out of range for this route, so for longer videos cut the clip and the words together.
| MAI-style field | Sume field | Conversion |
|---|---|---|
| text | text | Copy, drop empty strings |
| offsetMilliseconds | start | Divide by 1000 |
| offsetMilliseconds + durationMilliseconds | end | Divide by 1000, clamp to 60 |
| Words starting at 60 s or later | Dropped | Out of range |
| More than 1,200 words | Truncated | Split the clip |
The converter
The function was run on a two-word sample before publishing. It drops words that start after the 60-second ceiling and clamps the end of a word that straddles it.
def to_sume_words(mai, limit=60.0):
out = []
for phrase in mai.get("phrases", []):
for w in phrase.get("words", []):
text = w.get("text", "").strip()
start = w["offsetMilliseconds"] / 1000
end = start + w["durationMilliseconds"] / 1000
if text and start < limit:
out.append({"text": text[:200],
"start": round(start, 3),
"end": round(min(end, limit), 3)})
return out[:1200]
sample = {"phrases": [{"words": [
{"text": "Hello", "offsetMilliseconds": 120, "durationMilliseconds": 300},
{"text": "world", "offsetMilliseconds": 59800, "durationMilliseconds": 600}]}]}
print(to_sume_words(sample))Send it
Post the list with the video URL. The video must be a fetchable public HTTPS URL. Poll the job with GET /v1/jobs/{id}/status as with any Sume job. Use this path when your transcript is already corrected; use script_text instead when you only want the wording fixed and are happy with Sume's timing.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"video_url": "https://example.com/clip.mp4",
"words": [{"text": "Hello", "start": 0.12, "end": 0.42}]}'Edge cases that bite
Check three things after conversion. First, order: the words array should be sorted by start, and if your source phrases overlap (two speakers talking), sort before sending. Second, empty or whitespace-only tokens: drop them, because text must be at least one character. Third, punctuation: if the transcriber attaches punctuation to words, keep it, because captions read better with it, and if it returns it as separate tokens, merge it into the previous word.
If the video is longer than 60 seconds, split it into clips and shift each clip's word times back by that clip's start. A word that starts before a boundary and ends after it belongs to the earlier clip, which is why the converter clamps the end time instead of dropping the word.
Sources
Related posts
More in Developers
- Dart: submit a 30-second Wan 3.0 clip with package:http
Server-side Dart with package:http: POST wan-3.0 for 30 seconds, poll until completed and write the MP4. Keep the key out of the Flutter app.
- Deno 2.9.7 traceparent fix: tag a Sume job id in a Deno.serve log
Deno 2.9.7 extracts traceparent from Deno.serve regardless of header case. In a Sume webhook handler, log the job_id with the trace so a render is traceable.
- Deno: a 30-second Wan 3.0 video with fetch and top-level await
A short Deno script that submits a 30-second wan-3.0 job, polls to completion and writes clip.mp4, using only fetch and no npm packages.
- Deploy check: does your video model id and duration exist on Sume?
A Python preflight that reads GET /v1/videos/models and fails the deploy when VIDEO_MODEL is not a catalog id or the duration is outside supported_durations.
Written by Sume