Sume TTS word timestamps to caption cues for a narrated 60-second clip

Ask Sume TTS for timestamps.words, group them into cues and send them to video-captions so no recognition runs. About 25 cents for a 60-second narration.

6 min readSume
All posts

Set timestamps: {"words": true} on the TTS request and the completed job returns a words array with start and end seconds for each spoken word. Group those words into short cues and send them as cues to POST /v1/video-captions, and no speech recognition runs on the narration at all. For a one-minute narration the total is about 25 cents: roughly 5 cents of TTS plus $0.20 for captions.

The TTS figure assumes about 900 characters for a minute of narration (an assumption: 150 words a minute at six characters a word), which is 4.3 cents and bills as 5 cents.

Why use TTS timings

Speech recognition can mis-hear a brand name. The words you passed to TTS are the words that were spoken, so timings from the synthesis step carry your exact spelling. TTS results are monotonic, and with segmentation also set you can request gapless sentence segments, which require timestamps.words: true.

60-second narrated clip, Sume cost (read 2026-10-08)
StepRouteCost
Narration, about 900 charactersPOST /v1/tts-1.0/generate$0.05
Captions from cuesPOST /v1/video-captions$0.20
Total$0.25

Group words into cues

Cues take text up to 400 characters, with start and end between 0 and 60 seconds, at most 200 cues. Seven words per cue reads well on a phone. The key that holds the word text in the TTS result is not spelled out in the contract excerpt I read, so the function accepts either word or text.

def to_cues(words, per_cue=7):
    cues = []
    for i in range(0, len(words), per_cue):
        chunk = words[i:i + per_cue]
        label = lambda w: w.get("word") or w.get("text", "")
        cues.append({"text": " ".join(label(w) for w in chunk),
                     "start": chunk[0]["start"],
                     "end": min(chunk[-1]["end"], 60.0)})
    return cues[:200]

words = [{"word": t, "start": i * 0.4, "end": i * 0.4 + 0.35}
         for i, t in enumerate("New season, new fit, same low price today".split())]
print(to_cues(words, 4))

Mind the order

Cues are mutually exclusive with words, segments and script_text. The cue timings are relative to the narration audio, so they only line up if the narration starts at second 0 of the video. If you mix the narration into a timeline with an intro, shift every cue by the intro length first.

When to prefer speech recognition anyway

Cues from TTS timings describe the narration you generated. If the final video was re-edited, the voice was sped up, or a human re-recorded a line, those timings no longer match, and you are better off letting the caption route listen to the actual audio, using script_text to keep your spelling. Use TTS timings when the narration goes into the video untouched and you want exact wording with no recognition step.

Also keep clips within 60 seconds: both start and end in a cue are limited to that range. For longer narration, split the script into parts, generate each part as its own TTS job and caption each part as its own clip.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume