Captions for a voiceover: reuse TTS word timings or run STT?

If you generated the voice on Sume, ask for timestamps.words and skip a 1-cent-a-minute STT pass. If the audio is someone else's, run STT. Here is the split.

4 min readSume
All posts

If you generate the voiceover with Sume TTS, ask for word timestamps in the same request and use those timings for captions, which avoids a second transcription. If the audio was recorded by a person, you have no timings, so you run STT at $0.01 a minute. Both paths end at the same place: cues you burn into the video.

Two sources of timing

TTS accepts timestamps with words set to true, and segmentation by sentence requires it. STT returns the words it heard and, with segmentation, sentence ranges. The difference is where the text comes from: with TTS the text is your script, so nothing is misheard; with STT the text is a model's guess at the audio and needs a read-through.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
SituationTiming sourceExtra costText accuracy
You wrote the script and used Sume TTStimestamps.words from the TTS jobNone beyond the TTS jobExactly your script
A person recorded the voiceSTT job$0.01 per audio minute, 10 minutes max per jobCheck for misheard names
Recorded voice, but you have the scriptSTT for timing, your script for text$0.01 per minuteYour words, aligned to heard timing

The script-text shortcut

When you have both audio and a script, use the script as the caption text and the transcript only for timing. That keeps your spelling, brand names and numbers. Mixed cases are common in practice, because a human reads a script and makes small ad-libs, so compare before you publish.

  • Generated voice: request timestamps.words, keep the job result.
  • Recorded voice with no script: run STT, then correct the text.
  • Recorded voice with a script: align your script to the timing.

Which to pick today

For a weekly show where your team writes the script and generates the voice, keep the job results and use their timings: the text is already right. For interviews and customer calls, run STT and plan a correction pass. For a human-read script, align the script to the recorded timing and keep your own text. Whichever path you choose, burn captions from cues so the style stays consistent across clips.

A cost check

A 3-minute recorded voiceover costs about 3 cents to transcribe at list. That is small enough that the deciding factor is accuracy and editing time, not price. For synthesized voice, the timings arrive with the job you were already paying for.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume