AI voiceover for Shorts in order: TTS, then captions, then timeline
Best order for an AI voiceover Short: TTS with word timestamps, captions from those words, then a timeline render. Documented Sume limits and prices.

Make the voiceover first, with word timestamps; caption it second, feeding those words so nothing is transcribed again; render the timeline last. That order uses one TTS job, one $0.20 caption pass for clips up to 60 seconds, and one timeline render at $0.10 per output minute rounded up.
Reversing the order, for example captioning before you are sure of the script, means paying for steps twice.
Step 1: the voiceover
Send the script to TTS 1.0 or the TTS Router with timestamps.words set to true. The transcript limit is 20,000 characters and audio longer than 1200 seconds fails with tts_duration_exceeded, which is far beyond a Short. You get word start and end times in the result alongside the audio. Set the language explicitly for anything but English.
Listen before you go on. Changing a word later means rerunning this step and everything after it.
Step 2: captions from the words
The video captions endpoint accepts words, cues or segments, which skip speech recognition, and costs $0.20 for videos up to 60 seconds. Give it the TTS words so the text is exactly your script. Pick a style, or omit it to get slam for Latin text. Keep phrasing sensible with max_words, max_chars and pause_seconds, and place captions with anchor_ratio so they clear the platform's bottom UI.
Estimated duration over 60 seconds is rejected, so split longer videos.
Step 3: the timeline render
Timeline 1.0 takes the voiceover as audio.url, an optional soundtrack with gain_db, loop and a fade_out_seconds of up to 10, and output fades of 0 to 5 seconds. Ducking the music under the voice uses duck_db from 0 to 20 and needs a real spine. The render is billed at $0.10 per output minute rounded up.
| Step | Sume doc | Documented price |
|---|---|---|
| Voiceover | TTS contract | Not stated in this post |
| Captions | Video captions | $0.20 for videos up to 60 seconds |
| Join voice parts | Timeline audio concat | $0.01 flat per job |
| Final render | Timeline 1.0 | $0.10 per output minute, rounded up |
Disclosure
YouTube's guidance lists captions, voice cloning for your own voiceovers and AI-assisted scripts among things that do not need a disclosure, while realistic altered content does. Policies change, so read the page for your case before you publish. A Short that only uses an AI voice for your own script is the simple case the page describes.
Sources
Related posts
More in Use cases
- AI wedding invitation image: quote names and date so they render
Make an invitation card picture with an image API: quote every line, pick a portrait ratio, check the spelling, and add real print text in layout.
- AI whiteboard photo to a clean diagram: an image edit, step by step
Photograph the whiteboard, send it as a reference, and ask for a clean redraw. What an image model can and cannot keep, and how to check every label.
- AI wireframe to UI mockup: turn a hand sketch into a screen
Send a photo of your wireframe as a reference to GPT Image 2.5 on Sume, quote every label, and get a styled UI mockup back. Request, ratios and limits.
- AI workout music for fitness videos: tempo and intervals in a prompt
Make AI workout music by writing tempo as a number and the interval plan as timestamps. Sume has no BPM field, so here is how to prompt, check and trim a track.
Written by Sume