TTS word timings to karaoke captions: map words[] to text

Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.

5 min readSume
All posts

Ask Sume TTS for timestamps.words: true, read words[] from the finished job, and send the same entries to /v1/video-captions as words, renaming each token's text key to text. Sume then skips speech-to-text and burns exactly those words at exactly those times, so the caption cannot drift from the script you wrote.

This matters more now that fast voices exist. Microsoft's MAI-Voice-2.1-Flash is listed at 45 seconds of audio and 150 ms end-to-end latency (Microsoft AI, read 2026-10-05), and short clips tend to be captioned word by word.

The two shapes

The TTS request schema says that with timestamps.words true, the completed job result includes monotonic words[] with start and end seconds. The caption schema takes words as {text, start, end}, and the video captions guide says these are mutually exclusive with script_text, cues and segments. Send one wording source only.

A Python map

The code below renames the key and keeps only entries with a usable time. It accepts either word or text, because the caption side is strict and the TTS side is read from a job result.

Run it on the result you already fetched. It makes no network call, so you can test it on a saved result.

def to_caption_words(tts_words):
    out = []
    for w in tts_words:
        text = w.get("word") or w.get("text")
        start, end = w.get("start"), w.get("end")
        if not text or start is None or end is None:
            continue
        out.append({"text": text, "start": float(start), "end": float(end)})
    return out

def caption_body(video_url, tts_words, style="punch"):
    return {"video_url": video_url, "style": style,
            "words": to_caption_words(tts_words)}

if __name__ == "__main__":
    sample = [{"word": "Hello", "start": 0.0, "end": 0.4},
              {"word": "there", "start": 0.4, "end": 0.8}]
    print(caption_body("https://media.sume.com/artifacts/demo/clip.mp4", sample))

What to check

The video must carry the same audio the words came from, or the timings will not line up. Put the TTS file under the video with a Timeline 1.0 render, then caption the rendered MP4.

A standalone caption job is $0.20 for a video of up to 60 seconds under the current fixed estimate, and a restyle is billed as a render too. Check GET /v1/catalog for the live price.

Style rules still apply

Latin text works on slam, punch and tiktok-green. Korean text on those styles returns caption_hangul_text_latin_style, so pick a Hangul style for Korean words. Hangul styles such as black-outline, weight-shift and clip-wipe are the safe ones for Korean speech, and font is only accepted with them.

punch and tiktok-green do not support design overrides, while the other styles take a design object that changes colors, typography, placement, phrasing and motion for one request. Wrong values fail with a 400 at request time, so you do not pay for a bad render. To try a second look, send source_caption_id with a new style and the same words: the timings are reused and no second transcription runs.

The result is a captioned MP4 that burned your wording at TTS times, with no recognizer in between.

Words or cues

If you want words to appear in groups, send cues instead of words. A cue is one on-screen span with text, start and end, and it may hold a newline for a two-line card. You can build cues from segments[] of the same TTS job, which carry sentence text and gapless times. Word-level words give the karaoke look, and phrase-level cues give a quieter subtitle look. They are mutually exclusive, so choose before you submit.

Either way the wording comes from the script, so a brand spelling or a product name stays exactly as you typed it.

Keep the audio and the words paired in your records: the job id of the TTS take, the id of the caption job, and the style used. If a caption looks off later, you can tell whether the words, the timings or the render caused it.

Compare that with the other path. Without words, a caption job runs speech-to-text on the audio and burns what a recognizer heard, or aligns your script_text onto those timings and can fail with script_alignment_mismatch. Sending the TTS timings removes that failure mode, because there is no alignment step to fail. The trade is that you are trusting the TTS timings, so spot-check the first and last words against the audio before a batch run.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume