Burn captions from TTS word timings: no second transcription
Ask Sume TTS for word timings, group them into phrases and send them as caption cues. The caption job skips speech-to-text and still costs $0.20.

You can skip speech-to-text for an AI voiceover: request timestamps.words on the TTS job, group the word timings into phrases, and send them to the caption job as cues. Sume then burns your text at those times without running transcription, at the same $0.20 caption price for a video up to 60 seconds.
Why this beats transcribing your own voiceover
Transcription of a synthetic voice is a round trip: you wrote the script, the engine spoke it, and a recogniser guesses it back, sometimes wrongly. A product name that the recogniser mishears becomes a typo on screen. Sume's caption docs say that with words, cues or segments, Sume does not run speech-to-text and burns the text you sent at the times you sent.
Microsoft's MAI-Transcribe-2 page promises timestamps and domain biasing for transcription (read 2026-10-05). Those are good features for audio you did not write. For audio you did write, the timings from the generator are a better source.
The shape of the request
The TTS job takes timestamps: { words: true } and returns monotonic words[] with start and end in seconds on the completed result. The caption job takes cues as an array of text, start and end in seconds, and you may send only one of script_text, words, cues and segments.
The loop below splits the script on spaces, checks that the count equals the number of timed words, and groups four words per cue. It uses only the start and end of each timing, which are the fields the docs guarantee. If the counts differ, it stops instead of guessing.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def cues_from(script, timings, size=4):
tokens = script.split()
if len(tokens) != len(timings):
raise SystemExit("token and timing counts differ")
out = []
for i in range(0, len(tokens), size):
chunk = timings[i:i + size]
out.append({"text": " ".join(tokens[i:i + size]),
"start": chunk[0]["start"],
"end": chunk[-1]["end"]})
return out
def caption(video_url, cues, key):
r = requests.post("https://api.sume.com/v1/video-captions",
headers={**H, "Idempotency-Key": key},
json={"video_url": video_url, "cues": cues}, timeout=60)
r.raise_for_status()
return r.json()What changes and what does not
The price does not move, so the reason to choose cues is control and certainty, not savings. A failure from a mismatch between script and speech cannot happen when no recogniser is involved.
| Path | Speech-to-text runs | Text on screen | Price for a clip up to 60 s |
|---|---|---|---|
script_text | Yes, for timings | Your script, aligned to recognised words; can fail with script_alignment_mismatch | $0.20 |
cues from TTS timings | No | Exactly the text you sent | $0.20 |
| No text, no timings | Yes | Whatever is recognised | $0.20 |
Two caveats
- Splitting on spaces assumes one timed word per token. A script with hyphenated names or numbers may not line up, which is why the code stops on a count mismatch. Fix the tokenisation instead of loosening the check.
- Cue timings follow the audio you timed. If you later join parts or put the voice on a timeline, shift the cues by the same offset, or the text will land early.
Style and language
Cue captions use the same styles as any caption job: slam for Latin text by default, with Hangul styles for Korean. Because speech-to-text is out of the loop, the language field has nothing to hint, so leave it off.
A worked example
Take a 45-second voiceover of 112 words. With four words per cue the loop produces 28 cues, each on screen for roughly a second and a half, which suits a vertical video where lines should be short. Make the group size a parameter and look at the first render before you commit to a house style: three words feels punchy, six feels like a subtitle.
Then render once. The job is a single $0.20 caption render, and the timings were free because they came back with the audio you had already bought. If a client asks for a wording change after approval, edit the text only, keep the timings, and send the new cues: no second voice job is needed unless the spoken words changed too.
When to prefer script_text
Use script_text when the audio was not generated by Sume, for example a recorded narrator, because then you have no timings of your own and the recogniser's are the source of truth. Use cues when you made the audio and have the timings in hand.
Sources
Related posts
More in Developers
- Callback or polling for Omni 4K jobs on the Sume video API
Use callback_url on /v1/videos or mode webhook on the Video Router for Omni 4K batches: one signed request per job, not repeated polls. Checks included.
- Can I put my Sume API key in the MCP server URL? No, use a header
Do not put a Sume key in the MCP URL. The MCP auth spec bans tokens in the query string. Send a Bearer or x-api-key header, or use OAuth, in each client.
- Cancel a video chain midway: 409 job_generation_already_started
Cancel works only before generation starts. In a render, trim and captions chain, cancel queued jobs, let started jobs finish, and stop submitting steps.
- Cancel a Wan 3.0 job before it starts, and what a 1080p clip reserves
A Wan 3.0 job can be canceled only before generation starts; after that you get 409 job_generation_already_started. The reserve is $7.50 for 30 s at 1080p.
Written by Sume