TTS word timings to karaoke captions: map words[] to text
Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.

Ask Sume TTS for timestamps.words: true, read words[] from the finished job, and send the same entries to /v1/video-captions as words, renaming each token's text key to text. Sume then skips speech-to-text and burns exactly those words at exactly those times, so the caption cannot drift from the script you wrote.
This matters more now that fast voices exist. Microsoft's MAI-Voice-2.1-Flash is listed at 45 seconds of audio and 150 ms end-to-end latency (Microsoft AI, read 2026-10-05), and short clips tend to be captioned word by word.
The two shapes
The TTS request schema says that with timestamps.words true, the completed job result includes monotonic words[] with start and end seconds. The caption schema takes words as {text, start, end}, and the video captions guide says these are mutually exclusive with script_text, cues and segments. Send one wording source only.
A Python map
The code below renames the key and keeps only entries with a usable time. It accepts either word or text, because the caption side is strict and the TTS side is read from a job result.
Run it on the result you already fetched. It makes no network call, so you can test it on a saved result.
def to_caption_words(tts_words):
out = []
for w in tts_words:
text = w.get("word") or w.get("text")
start, end = w.get("start"), w.get("end")
if not text or start is None or end is None:
continue
out.append({"text": text, "start": float(start), "end": float(end)})
return out
def caption_body(video_url, tts_words, style="punch"):
return {"video_url": video_url, "style": style,
"words": to_caption_words(tts_words)}
if __name__ == "__main__":
sample = [{"word": "Hello", "start": 0.0, "end": 0.4},
{"word": "there", "start": 0.4, "end": 0.8}]
print(caption_body("https://media.sume.com/artifacts/demo/clip.mp4", sample))What to check
The video must carry the same audio the words came from, or the timings will not line up. Put the TTS file under the video with a Timeline 1.0 render, then caption the rendered MP4.
A standalone caption job is $0.20 for a video of up to 60 seconds under the current fixed estimate, and a restyle is billed as a render too. Check GET /v1/catalog for the live price.
Style rules still apply
Latin text works on slam, punch and tiktok-green. Korean text on those styles returns caption_hangul_text_latin_style, so pick a Hangul style for Korean words. Hangul styles such as black-outline, weight-shift and clip-wipe are the safe ones for Korean speech, and font is only accepted with them.
punch and tiktok-green do not support design overrides, while the other styles take a design object that changes colors, typography, placement, phrasing and motion for one request. Wrong values fail with a 400 at request time, so you do not pay for a bad render. To try a second look, send source_caption_id with a new style and the same words: the timings are reused and no second transcription runs.
The result is a captioned MP4 that burned your wording at TTS times, with no recognizer in between.
Words or cues
If you want words to appear in groups, send cues instead of words. A cue is one on-screen span with text, start and end, and it may hold a newline for a two-line card. You can build cues from segments[] of the same TTS job, which carry sentence text and gapless times. Word-level words give the karaoke look, and phrase-level cues give a quieter subtitle look. They are mutually exclusive, so choose before you submit.
Either way the wording comes from the script, so a brand spelling or a product name stays exactly as you typed it.
Keep the audio and the words paired in your records: the job id of the TTS take, the id of the caption job, and the style used. If a caption looks off later, you can tell whether the words, the timings or the render caused it.
Compare that with the other path. Without words, a caption job runs speech-to-text on the audio and burns what a recognizer heard, or aligns your script_text onto those timings and can fail with script_alignment_mismatch. Sending the TTS timings removes that failure mode, because there is no alignment step to fail. The trade is that you are trusting the TTS timings, so spot-check the first and last words against the audio before a batch run.
Sources
Related posts
More in Developers
- Turn STT words into paragraphs: break at pauses over 1.2 seconds
Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.
- AI music API in TypeScript: generate, wait and save an MP3
Call Sume's Music Router from TypeScript: generateMusicRouter, waitForJob, then download the audio artifact. $0.125 per track, Google lists Lyria 3.5 at $0.08.
- TypeScript exhaustive switch over a Sume run's terminal status
A run ends as completed, failed, canceled or skipped, and the last two send no webhook. Use a never check so a new status fails the build.
- TypeScript webhook verifier for Wan 3.0 clips: refuse an empty secret
A TypeScript verifier for Sume webhooks: HMAC SHA 256 over timestamp.body, rotation entries, a 5 minute replay window and a hard refusal of an empty secret.
Written by Sume