Reuse TTS word timestamps as caption words, skip a second STT

You already know what the voice said and when. Feed the TTS word timings to the caption job as `words` so brand names are never misheard by speech-to-text.

6 min readSume
All posts

If an AI voice read your script, do you need speech-to-text to caption the video? No. The TTS job can return the word timings itself, and the caption job accepts words and skips transcription entirely. You burn exactly your script's spelling at the times the voice actually spoke it.

This matters more as voices get better. ElevenLabs says Eleven v4 (read 2026-10-04) is more expressive across more than 90 languages; expressive delivery tends to change pacing, and pacing is what a fixed caption schedule gets wrong.

The three pieces

First, ask the TTS job for timings: timestamps: { words: true } adds a monotonic words[] to the finished result, each item with word, start and end in seconds. Second, put the voice into a video and host it so you have a public HTTPS video_url. Third, send the words to POST /v1/video-captions.

The caption job names its fields text, start, end, so rename word to text. words is mutually exclusive with script_text, cues and segments. Sending words skips speech-to-text and burns exactly that copy at those times.

Convert and submit

This sketch assumes you already have the TTS result's words list and the hosted video URL. If the voice starts later than zero in the video, add that offset to every start and end first.

import os, requests

tts_words = [
    {"word": "Meet", "start": 0.00, "end": 0.28},
    {"word": "Sume", "start": 0.30, "end": 0.71},
]
offset = 0.0
words = [{"text": w["word"], "start": w["start"] + offset,
          "end": w["end"] + offset} for w in tts_words]
r = requests.post(
    "https://api.sume.com/v1/video-captions",
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Idempotency-Key": "caption-from-tts-001"},
    json={"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
          "style": "slam", "words": words})
print(r.status_code)

When to use which input

Use words when the audio is entirely synthetic and you hold the timings. Use script_text when the audio is a recording and you want STT timing with your wording; it can fail with script_alignment_mismatch or script_alignment_failed. Use cues or segments for phrase-level overlay text, including silent clips.

The TTS timings describe the TTS audio, not the final video. If you trimmed, padded or re-timed the voice in a render, re-base the timings first. The words array shares one schema with STT results, so the same conversion works for a transcript you edited by hand.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume