TTS word timings straight into captions: a 45-second ad for 33 cents
Ask Sume TTS for timestamps.words and send them as words on the caption job: no speech-to-text. 700 characters, render and captions come to $0.33.

For a voice-over that Sume TTS made, you already hold word timings: request timestamps.words: true, then send those words to the caption job as words with text, start and end, and Sume skips speech-to-text and burns exactly those words at exactly those times. A 45-second ad with 700 characters of script costs $0.03325 for the voice, $0.10 for the one-minute render and $0.20 for captions: $0.33325.
Sources: the Sume TTS request schema, the video captions docs and request contract, and the API catalog (read 2026-10-09).
The mapping
TTS returns words[] with start and end seconds from the audio start. The caption contract wants objects of text, start, end, with 1 to 1,200 items, text up to 200 characters, and start and end between 0 and 60 seconds. Map each word entry to text and keep its start and end times. You can send only one of script_text, words, cues and segments.
The offsets only line up if the voice starts at 0 in the final video. If you trimmed with source_in or joined files first, add the matching offset to each time before you send them.
| Step | Detail | Cost |
|---|---|---|
| TTS with timestamps.words | 700 characters x $0.0475 per 1,000 | $0.03325 |
| Timeline render | 45 s rounds up to 1 minute | $0.10 |
| Caption job with words | Fixed, videos up to 60 s | $0.20 |
| Total | $0.33325 |
Why do it
The captions show your script exactly, with the spelling of product names as you wrote them, and no second transcription can mishear a brand. The price is the same as a speech-to-captions job; the gain is accuracy and one less variable. With script_text, Sume aligns your wording to its own speech-to-text timings and can fail alignment; with words there is nothing to align.
The cost is that your word times must be right. Check the first and last word against the audio before you burn a batch.
Silent edits and music
If the final video has music over the voice, the words still come from the TTS file, so a loud bed cannot confuse the timings. Use soundtrack.duck_db in the timeline to keep the voice clear. For a clip with no voice at all, use cues instead, which is the silent-clip path.
A small check before the burn
Compare the number of words in your script with the length of the words[] array. If they differ a lot, punctuation or numbers may have been split or merged differently from how you expect. Fix the mismatch in your mapping, not in the caption job.
Because the caption contract allows up to 1,200 words and a 60-second window, a 45-second ad at about 2.5 words per second is around 112 words, far under the limit.
Both calls take an Idempotency-Key. Submit the render first, wait for its video_url, then submit the caption job with that URL and your words, since the caption job needs a public HTTPS video it can fetch.
Sources
Related posts
More in Developers
- 20 Nano Banana 2.1 2K images: $3.00 and one jobs_result read
Twenty Nano Banana 2.1 images at 2K cost $3.00 on Sume. Over hosted MCP, one jobs_result call with 20 job_ids reads them back; failed ids are named.
- 20 Omni Flash 1.1 clips in one jobs_wait wave: $12.50 at 720p
Twenty 5-second Gemini Omni Flash 1.1 clips at 720p cost $12.50 on Sume ($0.625 each) and fit one jobs_wait call of 20 ids. Cost table and the call body.
- TypeScript: check 8:1 against /v1/images/models before you POST
A 20-line TypeScript guard that reads the aspect_ratio values for a Sume image model and refuses a ratio it does not list, instead of a 400 at request time.
- usage.cost on a Sume video poll is nullable: guard it before you sum
The /v1/videos poll schema allows usage.cost to be null. Sum only completed jobs, keep unknowns apart in a ledger, and compare to the Wan 3.0 $3.75 estimate.
Written by Sume