Caption generator from audio: STT words into video captions, $0.21

Have a recording and a separate video? Transcribe on Sume at $0.01 a minute, then burn the word timings with a $0.20 caption job. Mapping code included.

5 min readSume
All posts

"Caption generator from audio" is on Google's autocomplete (read 2026-10-07). Sume's video captions job is built around a video: it transcribes the clip's own speech or aligns a script. When your words live in a separate recording, say a voice memo laid over screen footage, you can do the speech step yourself and hand the result to the caption job.

The caption job takes words with text, start and end in seconds, and with words Sume does not run speech to text. You can send only one of script_text, words, cues and segments.

Two calls

Total for a one-minute clip is $0.21. The only code you write is the rename, shown here with sample timings so it runs as is.

  • POST /v1/stt-1.0/transcribe on the audio: $0.01 per audio minute, up to 10 minutes. The result has words[] as {word, start, end}.
  • POST /v1/video-captions with the video's public HTTPS video_url and those words renamed to {text, start, end}: $0.20 for a video up to 60 seconds.
import json

stt_words = [
    {"word": "Say", "start": 0.0, "end": 0.3},
    {"word": "hello", "start": 0.3, "end": 0.7},
]

payload = {
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "punch",
    "words": [
        {"text": w["word"], "start": w["start"], "end": w["end"]}
        for w in stt_words
    ],
}
print(json.dumps(payload, indent=2))

Alignment limits and the alternative

If the audio is already in the video, skip all of this: send the video_url and let the caption job do its own speech step at the same $0.20.

Caption inputs, read 2026-10-07
InputWhat Sume doesFailure code
video_url onlyRuns STT on the clip's audio, burns the wordscaption_no_speech if the clip is silent
video_url + script_textKeeps STT timings, aligns your text to themscript_alignment_mismatch, script_alignment_failed
video_url + wordsBurns your words at your times, no STTA 400 for out-of-range style fields
video_url + cues or segmentsBurns phrase cards at your times, no STTAs above

Language and style

Pass the same language_code hint to STT that you would pass as language to captions (ko, en). For Korean text pick a Hangul style such as black-outline or korean-ad; the Latin styles reject Korean text with a 400. The language field never selects the style.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume