Caption generator from audio: STT words into video captions, $0.21
Have a recording and a separate video? Transcribe on Sume at $0.01 a minute, then burn the word timings with a $0.20 caption job. Mapping code included.

"Caption generator from audio" is on Google's autocomplete (read 2026-10-07). Sume's video captions job is built around a video: it transcribes the clip's own speech or aligns a script. When your words live in a separate recording, say a voice memo laid over screen footage, you can do the speech step yourself and hand the result to the caption job.
The caption job takes words with text, start and end in seconds, and with words Sume does not run speech to text. You can send only one of script_text, words, cues and segments.
Two calls
Total for a one-minute clip is $0.21. The only code you write is the rename, shown here with sample timings so it runs as is.
POST /v1/stt-1.0/transcribeon the audio: $0.01 per audio minute, up to 10 minutes. The result haswords[]as{word, start, end}.POST /v1/video-captionswith the video's public HTTPSvideo_urland those words renamed to{text, start, end}: $0.20 for a video up to 60 seconds.
import json
stt_words = [
{"word": "Say", "start": 0.0, "end": 0.3},
{"word": "hello", "start": 0.3, "end": 0.7},
]
payload = {
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "punch",
"words": [
{"text": w["word"], "start": w["start"], "end": w["end"]}
for w in stt_words
],
}
print(json.dumps(payload, indent=2))Alignment limits and the alternative
If the audio is already in the video, skip all of this: send the video_url and let the caption job do its own speech step at the same $0.20.
| Input | What Sume does | Failure code |
|---|---|---|
video_url only | Runs STT on the clip's audio, burns the words | caption_no_speech if the clip is silent |
video_url + script_text | Keeps STT timings, aligns your text to them | script_alignment_mismatch, script_alignment_failed |
video_url + words | Burns your words at your times, no STT | A 400 for out-of-range style fields |
video_url + cues or segments | Burns phrase cards at your times, no STT | As above |
Language and style
Pass the same language_code hint to STT that you would pass as language to captions (ko, en). For Korean text pick a Hangul style such as black-outline or korean-ad; the Latin styles reject Korean text with a 400. The language field never selects the style.
Sources
Related posts
More in Developers
- Captions out of sync with the audio: check STT word times and offsets
Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.
- Chapter timestamps for narrated audio from concat segment offsets
Join one TTS file per chapter with timeline audio concat, then turn the returned segments[] start offsets into mm:ss chapter lines with a short Python script.
- Check a reference image URL before sending it to Sume Images
Sume rejects localhost, private-network and non-HTTPS reference URLs before submission. A short Python pre-flight check that catches them first.
- Check an Omni edit kept the rest of the clip: video-frames pairs
Compare stills from the source and the edited clip at the same timestamps with video-frames on Sume. A script that submits both extracts, plus what to look for.
Written by Sume