Convert MAI-Transcribe-2 word offsets to Sume video caption words

Turn word timings in milliseconds into the seconds-based words array Sume video-captions accepts. A short Python converter plus the 60 s and 1,200-word limits.

6 min readSume
All posts

If you transcribe with MAI-Transcribe-2 and want the words burned onto a short video, convert each word's millisecond offset and duration into start and end in seconds and send the list as words to POST /v1/video-captions. When words is present Sume skips its own speech-to-text and burns exactly those words at exactly those times. The job costs $0.20 for a clip up to 60 seconds.

This post assumes the Azure Speech batch response shape, in which each phrase has a words list and every word has text, offsetMilliseconds and durationMilliseconds. Check your own response before relying on those names; Microsoft's launch post for MAI-Transcribe-2 (read 2026-10-08) does not document the field layout.

The limits you convert into

Sume's words entries take text (1 to 200 characters), and start and end as numbers from 0 to 60. The array holds at most 1,200 entries and cannot be combined with script_text, cues or segments. A clip over 60 seconds is out of range for this route, so for longer videos cut the clip and the words together.

Field mapping, MAI-style word to Sume caption word (read 2026-10-08)
MAI-style fieldSume fieldConversion
texttextCopy, drop empty strings
offsetMillisecondsstartDivide by 1000
offsetMilliseconds + durationMillisecondsendDivide by 1000, clamp to 60
Words starting at 60 s or laterDroppedOut of range
More than 1,200 wordsTruncatedSplit the clip

The converter

The function was run on a two-word sample before publishing. It drops words that start after the 60-second ceiling and clamps the end of a word that straddles it.

def to_sume_words(mai, limit=60.0):
    out = []
    for phrase in mai.get("phrases", []):
        for w in phrase.get("words", []):
            text = w.get("text", "").strip()
            start = w["offsetMilliseconds"] / 1000
            end = start + w["durationMilliseconds"] / 1000
            if text and start < limit:
                out.append({"text": text[:200],
                            "start": round(start, 3),
                            "end": round(min(end, limit), 3)})
    return out[:1200]

sample = {"phrases": [{"words": [
    {"text": "Hello", "offsetMilliseconds": 120, "durationMilliseconds": 300},
    {"text": "world", "offsetMilliseconds": 59800, "durationMilliseconds": 600}]}]}
print(to_sume_words(sample))

Send it

Post the list with the video URL. The video must be a fetchable public HTTPS URL. Poll the job with GET /v1/jobs/{id}/status as with any Sume job. Use this path when your transcript is already corrected; use script_text instead when you only want the wording fixed and are happy with Sume's timing.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video_url": "https://example.com/clip.mp4",
       "words": [{"text": "Hello", "start": 0.12, "end": 0.42}]}'

Edge cases that bite

Check three things after conversion. First, order: the words array should be sorted by start, and if your source phrases overlap (two speakers talking), sort before sending. Second, empty or whitespace-only tokens: drop them, because text must be at least one character. Third, punctuation: if the transcriber attaches punctuation to words, keep it, because captions read better with it, and if it returns it as separate tokens, merge it into the previous word.

If the video is longer than 60 seconds, split it into clips and shift each clip's word times back by that clip's start. A word that starts before a boundary and ends after it belongs to the earlier clip, which is why the converter clamps the end time instead of dropping the word.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume