Silent captioned Reel: drop the audio first, burn with words
To burn captions on a Reel with no sound, run STT first, trim with audio drop, then send your own words: $0.02 + $0.01 + $0.20 = $0.23.

A silent captioned Reel needs the words before the sound is removed, because speech-to-text on a clip without audio fails. Run STT on the source ($0.01), cut the Reel with audio: "drop" ($0.02), then pass the shifted words to video captions as words ($0.20): $0.23 in all.
Why the order matters
Video captions run speech-to-text unless you send script_text, words, cues or segments. If the clip has no audible speech and you send none of them, the job fails with caption_no_speech and next_action: use_overlay_captions. A video inspect with transcribe: true on a clip with no audio track fails with inspect_source_has_no_audio. So the words must come from the original, before the audio is dropped.
| Step | Order | Rate |
|---|---|---|
| Video inspect with transcribe (1 minute) | 1 | $0.01 |
| Video trim with audio drop | 2 | $0.02 |
| Video captions with words | 3 | $0.20 |
| Total | $0.23 |
Steps
Re-base the words against the trim's actual_start_seconds (see re-basing), keep them inside 60 s, and send only words, because the request accepts one of script_text, words, cues and segments.
trim_body = {"video_url": SRC, "start": 1.69, "end": 27.92, "audio": "drop"}
caption_body = {
"video_url": TRIMMED_URL,
"style": "slam",
"words": [{"text": "Meet", "start": 0.0, "end": 0.31},
{"text": "the", "start": 0.31, "end": 0.42}],
}Gotchas
- Word entries use
text,startandendin seconds; the same field names are used forcuesandsegments. - Silent output has no spine for Timeline; if you then join clips, the render needs its own audio or an explicit silent choice.
- Latin styles cannot render Korean; Korean text on
slam,punchortiktok-greenreturnscaption_hangul_text_latin_style. - Caption timings are capped at 60 s.
Sources
Related posts
More in Use cases
- Six Nano Banana 2.1 start frames, six Wan 3.0 clips: a 3-min Short
Six 9:16 Nano Banana 2.1 stills ($0.60 at 1K) feed six 30 s Wan 3.0 jobs as first frames: $23.40 at 720p for a 180-second Short, with Timeline.
- Six podcast ad-read variants at 500 characters: TTS cost and tags
Six 500-character ad reads cost $0.1425 on Sume TTS ($0.02375 each). Tag each take with the metadata field and one Idempotency-Key per variant.
- A 60-second Short as twenty 3-second Omni cuts: $7.50 at 720p
Twenty 3-second Omni Flash 1.1 clips at 720p cost $7.50 and at 1080p $11.25, the same as six 10 s clips. Twenty slots pass 12, so Timeline chunks the render.
- Spelling bee: 200 word-pronunciation clips on Sume TTS, one job each
200 words of about 12 characters each cost $0.1140 on Sume TTS as 200 jobs, $0.000570 each. One 2,400-character job is $0.1140 but gives one file.
Written by Sume