Captions for a voiceover: reuse TTS word timings or run STT?
If you generated the voice on Sume, ask for timestamps.words and skip a 1-cent-a-minute STT pass. If the audio is someone else's, run STT. Here is the split.

If you generate the voiceover with Sume TTS, ask for word timestamps in the same request and use those timings for captions, which avoids a second transcription. If the audio was recorded by a person, you have no timings, so you run STT at $0.01 a minute. Both paths end at the same place: cues you burn into the video.
Two sources of timing
TTS accepts timestamps with words set to true, and segmentation by sentence requires it. STT returns the words it heard and, with segmentation, sentence ranges. The difference is where the text comes from: with TTS the text is your script, so nothing is misheard; with STT the text is a model's guess at the audio and needs a read-through.
| Situation | Timing source | Extra cost | Text accuracy |
|---|---|---|---|
| You wrote the script and used Sume TTS | timestamps.words from the TTS job | None beyond the TTS job | Exactly your script |
| A person recorded the voice | STT job | $0.01 per audio minute, 10 minutes max per job | Check for misheard names |
| Recorded voice, but you have the script | STT for timing, your script for text | $0.01 per minute | Your words, aligned to heard timing |
The script-text shortcut
When you have both audio and a script, use the script as the caption text and the transcript only for timing. That keeps your spelling, brand names and numbers. Mixed cases are common in practice, because a human reads a script and makes small ad-libs, so compare before you publish.
- Generated voice: request timestamps.words, keep the job result.
- Recorded voice with no script: run STT, then correct the text.
- Recorded voice with a script: align your script to the timing.
Which to pick today
For a weekly show where your team writes the script and generates the voice, keep the job results and use their timings: the text is already right. For interviews and customer calls, run STT and plan a correction pass. For a human-read script, align the script to the recorded timing and keep your own text. Whichever path you choose, burn captions from cues so the style stays consistent across clips.
A cost check
A 3-minute recorded voiceover costs about 3 cents to transcribe at list. That is small enough that the deciding factor is accuracy and editing time, not price. For synthesized voice, the timings arrive with the job you were already paying for.
Sources
Related posts
More in Media tools
- Captions language field: a speech hint, not a style or font
On Sume video captions, language only hints speech-to-text. It never picks the style or font, so Korean needs a Hangul style such as black-outline.
- Check a Short has sound and picture before adding it to a YouTube show
YouTube asks shows for high-quality video and sound and says no-visuals videos likely will not count. Probe each Short with Sume video inspect first.
- compose_video_has_no_audio: a silent clip under a still on Sume
Timeline compose warns compose_video_has_no_audio and still completes. The output is silent. How to add a voice-over afterward with a Timeline render.
- Copy space in an AI image: leave room for text, overlay it in code
Ask for empty space in the prompt, check the region with Pillow, and put the headline on top in code, so the text is exact and the picture stays a picture.
Written by Sume