TTS word timings to burned-in captions: send them as words on Sume
Sume's TTS can return word start and end times; the caption job accepts words with text, start and end and skips transcription. How to wire them together.

You can caption a video made from a Sume voiceover without transcribing it again. Request the voiceover with timestamps: { words: true }, read the words[] timings off the finished job, and send them to the caption job as words, each with text, start and end in seconds. Sume's caption guide says words, cues and segments skip speech-to-text and burn exactly that copy at those times.
The two sides of the hand-off
On the TTS side, the schema says words: true returns monotonic words[] with start and end seconds on the job result. On the caption side, Video captions says authored timing skips transcription. Check the key names in your own job result before you map them, since the caption job wants text, start and end.
| Step | Field | Notes |
|---|---|---|
| TTS request | timestamps.words true | Returns words[] with start and end seconds |
| Caption request | words[] | text, start, end; end must be greater than start |
| Caption limits | Up to 1,200 words | Videos up to 60 seconds |
| Alternative | cues or segments | Phrase-level cards, same skip of transcription |
| Do not combine | script_text, words, cues, segments | Mutually exclusive |
Why bother
Captions made from your own timings carry the exact script wording, so a brand name that a transcriber would misspell stays correct, and you avoid paying a speech-to-text pass. You still pay for the render: a standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate.
Timings measure the voiceover alone. If you add a lead-in before the voice starts in the final video, add that offset to every start and end.
Sources
Related posts
More in Developers
- Validate video duration and resolution in Python before you submit
Fetch GET /v1/videos/models and check duration, resolution and aspect_ratio per model in about 25 lines of Python, before a Sume video job fails.
- Veo 2.0 and Veo 3.0 shut down June 30: what model id to call now
Google retired veo-2.0 and veo-3.0 ids on 2026-06-30. See which Veo and Omni ids the Gemini docs list now, and how to move the call to Sume.
- AI video API fallback: retry on another model when a job fails
Chain seedance-2.5, seedance-2 and kling-3 on Sume: poll status_url, read the job error category, and resubmit the brief to the next model.
- Caption inputs on Sume: script_text, words, cues or segments?
Four caption inputs and only one may be sent. When to use script_text with transcription, word timings, or phrase cues for a silent clip. Plus the errors.
Written by Sume