Sume TTS word timestamps to caption cues for a narrated 60-second clip
Ask Sume TTS for timestamps.words, group them into cues and send them to video-captions so no recognition runs. About 25 cents for a 60-second narration.

Set timestamps: {"words": true} on the TTS request and the completed job returns a words array with start and end seconds for each spoken word. Group those words into short cues and send them as cues to POST /v1/video-captions, and no speech recognition runs on the narration at all. For a one-minute narration the total is about 25 cents: roughly 5 cents of TTS plus $0.20 for captions.
The TTS figure assumes about 900 characters for a minute of narration (an assumption: 150 words a minute at six characters a word), which is 4.3 cents and bills as 5 cents.
Why use TTS timings
Speech recognition can mis-hear a brand name. The words you passed to TTS are the words that were spoken, so timings from the synthesis step carry your exact spelling. TTS results are monotonic, and with segmentation also set you can request gapless sentence segments, which require timestamps.words: true.
| Step | Route | Cost |
|---|---|---|
| Narration, about 900 characters | POST /v1/tts-1.0/generate | $0.05 |
| Captions from cues | POST /v1/video-captions | $0.20 |
| Total | $0.25 |
Group words into cues
Cues take text up to 400 characters, with start and end between 0 and 60 seconds, at most 200 cues. Seven words per cue reads well on a phone. The key that holds the word text in the TTS result is not spelled out in the contract excerpt I read, so the function accepts either word or text.
def to_cues(words, per_cue=7):
cues = []
for i in range(0, len(words), per_cue):
chunk = words[i:i + per_cue]
label = lambda w: w.get("word") or w.get("text", "")
cues.append({"text": " ".join(label(w) for w in chunk),
"start": chunk[0]["start"],
"end": min(chunk[-1]["end"], 60.0)})
return cues[:200]
words = [{"word": t, "start": i * 0.4, "end": i * 0.4 + 0.35}
for i, t in enumerate("New season, new fit, same low price today".split())]
print(to_cues(words, 4))Mind the order
Cues are mutually exclusive with words, segments and script_text. The cue timings are relative to the narration audio, so they only line up if the narration starts at second 0 of the video. If you mix the narration into a timeline with an intro, shift every cue by the intro length first.
When to prefer speech recognition anyway
Cues from TTS timings describe the narration you generated. If the final video was re-edited, the voice was sped up, or a human re-recorded a line, those timings no longer match, and you are better off letting the caption route listen to the actual audio, using script_text to keep your spelling. Use TTS timings when the narration goes into the video untouched and you want exact wording with no recognition step.
Also keep clips within 60 seconds: both start and end in a cue are limited to that range. For longer narration, split the script into parts, generate each part as its own TTS job and caption each part as its own clip.
Sources
Related posts
More in Developers
- Sume TypeScript SDK waitForJob: 20-minute timeout, job keeps billing
How @sume-com/sdk waitForJob polls a generation job, what SumeJobTimeoutError means, and why a client timeout does not cancel or refund the job.
- Webhook endpoint down: redeliver a Sume video job after the retries
Sume retries a job webhook 10 times, 30 seconds apart. If your receiver was down longer, POST /v1/jobs/{id}/webhook/redeliver re-sends the terminal payload.
- waitForJob times out at 20 minutes: why it does not fit a Vercel route
Sume SDK waitForJob waits 20 minutes by default and the job keeps billing if it throws. Vercel functions default to 300 s, so wait in a worker.
- Sume wave_size_hint and a Worker subrequest limit: submit in waves
A Worker fan-out of Sume jobs hits 50 subrequests on Free. Size each wave from generation_limits, not from the hint alone, and stop at queue_capacity_remaining.
Written by Sume