TTS word timings to burned-in captions: send them as words on Sume

Sume's TTS can return word start and end times; the caption job accepts words with text, start and end and skips transcription. How to wire them together.

5 min readSume
All posts

You can caption a video made from a Sume voiceover without transcribing it again. Request the voiceover with timestamps: { words: true }, read the words[] timings off the finished job, and send them to the caption job as words, each with text, start and end in seconds. Sume's caption guide says words, cues and segments skip speech-to-text and burn exactly that copy at those times.

The two sides of the hand-off

On the TTS side, the schema says words: true returns monotonic words[] with start and end seconds on the job result. On the caption side, Video captions says authored timing skips transcription. Check the key names in your own job result before you map them, since the caption job wants text, start and end.

Fields on each side of the hand-off, from Sume's tool schema and caption guide, read 2026-10-03
StepFieldNotes
TTS requesttimestamps.words trueReturns words[] with start and end seconds
Caption requestwords[]text, start, end; end must be greater than start
Caption limitsUp to 1,200 wordsVideos up to 60 seconds
Alternativecues or segmentsPhrase-level cards, same skip of transcription
Do not combinescript_text, words, cues, segmentsMutually exclusive

Why bother

Captions made from your own timings carry the exact script wording, so a brand name that a transcriber would misspell stays correct, and you avoid paying a speech-to-text pass. You still pay for the render: a standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate.

Timings measure the voiceover alone. If you add a lead-in before the voice starts in the final video, add that offset to every start and end.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume