TTS word timings straight into captions: a 45-second ad for 33 cents

Ask Sume TTS for timestamps.words and send them as words on the caption job: no speech-to-text. 700 characters, render and captions come to $0.33.

5 min readSume
All posts

For a voice-over that Sume TTS made, you already hold word timings: request timestamps.words: true, then send those words to the caption job as words with text, start and end, and Sume skips speech-to-text and burns exactly those words at exactly those times. A 45-second ad with 700 characters of script costs $0.03325 for the voice, $0.10 for the one-minute render and $0.20 for captions: $0.33325.

Sources: the Sume TTS request schema, the video captions docs and request contract, and the API catalog (read 2026-10-09).

The mapping

TTS returns words[] with start and end seconds from the audio start. The caption contract wants objects of text, start, end, with 1 to 1,200 items, text up to 200 characters, and start and end between 0 and 60 seconds. Map each word entry to text and keep its start and end times. You can send only one of script_text, words, cues and segments.

The offsets only line up if the voice starts at 0 in the final video. If you trimmed with source_in or joined files first, add the matching offset to each time before you send them.

45-second ad, voice to captions, as of 2026-10-09
StepDetailCost
TTS with timestamps.words700 characters x $0.0475 per 1,000$0.03325
Timeline render45 s rounds up to 1 minute$0.10
Caption job with wordsFixed, videos up to 60 s$0.20
Total$0.33325

Why do it

The captions show your script exactly, with the spelling of product names as you wrote them, and no second transcription can mishear a brand. The price is the same as a speech-to-captions job; the gain is accuracy and one less variable. With script_text, Sume aligns your wording to its own speech-to-text timings and can fail alignment; with words there is nothing to align.

The cost is that your word times must be right. Check the first and last word against the audio before you burn a batch.

Silent edits and music

If the final video has music over the voice, the words still come from the TTS file, so a loud bed cannot confuse the timings. Use soundtrack.duck_db in the timeline to keep the voice clear. For a clip with no voice at all, use cues instead, which is the silent-clip path.

A small check before the burn

Compare the number of words in your script with the length of the words[] array. If they differ a lot, punctuation or numbers may have been split or merged differently from how you expect. Fix the mismatch in your mapping, not in the caption job.

Because the caption contract allows up to 1,200 words and a 60-second window, a 45-second ad at about 2.5 words per second is around 112 words, far under the limit.

Both calls take an Idempotency-Key. Submit the render first, wait for its video_url, then submit the caption job with that URL and your words, since the caption job needs a public HTTPS video it can fetch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume