AI voiceover for Shorts in order: TTS, then captions, then timeline

Best order for an AI voiceover Short: TTS with word timestamps, captions from those words, then a timeline render. Documented Sume limits and prices.

4 min readSume
All posts

Make the voiceover first, with word timestamps; caption it second, feeding those words so nothing is transcribed again; render the timeline last. That order uses one TTS job, one $0.20 caption pass for clips up to 60 seconds, and one timeline render at $0.10 per output minute rounded up.

Reversing the order, for example captioning before you are sure of the script, means paying for steps twice.

Step 1: the voiceover

Send the script to TTS 1.0 or the TTS Router with timestamps.words set to true. The transcript limit is 20,000 characters and audio longer than 1200 seconds fails with tts_duration_exceeded, which is far beyond a Short. You get word start and end times in the result alongside the audio. Set the language explicitly for anything but English.

Listen before you go on. Changing a word later means rerunning this step and everything after it.

Step 2: captions from the words

The video captions endpoint accepts words, cues or segments, which skip speech recognition, and costs $0.20 for videos up to 60 seconds. Give it the TTS words so the text is exactly your script. Pick a style, or omit it to get slam for Latin text. Keep phrasing sensible with max_words, max_chars and pause_seconds, and place captions with anchor_ratio so they clear the platform's bottom UI.

Estimated duration over 60 seconds is rejected, so split longer videos.

Step 3: the timeline render

Timeline 1.0 takes the voiceover as audio.url, an optional soundtrack with gain_db, loop and a fade_out_seconds of up to 10, and output fades of 0 to 5 seconds. Ducking the music under the voice uses duck_db from 0 to 20 and needs a real spine. The render is billed at $0.10 per output minute rounded up.

Costs of each documented step (read 2026-10-03)
StepSume docDocumented price
VoiceoverTTS contractNot stated in this post
CaptionsVideo captions$0.20 for videos up to 60 seconds
Join voice partsTimeline audio concat$0.01 flat per job
Final renderTimeline 1.0$0.10 per output minute, rounded up

Disclosure

YouTube's guidance lists captions, voice cloning for your own voiceovers and AI-assisted scripts among things that do not need a disclosure, while realistic altered content does. Policies change, so read the page for your case before you publish. A Short that only uses an AI voice for your own script is the simple case the page describes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume