Reddit 15-second video ad: burn captions so it reads on mute
For a 15-second Reddit video ad, burn captions into the cut with POST /v1/video-captions: one $0.20 job, style slam for English, no new edit.

To make a 15-second Reddit video ad readable with the sound off, run the trimmed cut through POST /v1/video-captions, which burns captions into the picture and returns a new video. For English speech, leave style unset and Sume picks slam; the job is priced at $0.20 for clips up to 60 seconds. Reddit's 15-second Engaged Video Views option is listed as a public beta from 2026-09-01, with longer videos billing at 15 s (read 2026-10-06), so the captions only need to cover those 15 seconds.
Facts: Video captions: burn captions onto a video for the endpoint and price, and the Social media platform updates, October 2026 (read 2026-10-06) for the Reddit bid.
Why caption after the trim, not before?
Caption after you trim. A caption job transcribes the clip it receives, so a 15-second input is a 15-second job and nothing is burned for seconds that will not be bought. If you caption the 30-second master first and trim after, you pay the same $0.20 but you must also keep the caption timing aligned to the new start time. Order the steps trim, then captions.
Which options matter for a short ad?
Three fields do most of the work. style sets the look and motion; unset, Latin text resolves to slam and Korean text to black-outline. design overrides one request's colors, typography, placement and phrasing, for example phrasing.max_words to keep two or three words on screen. script_text makes the burned text match your script while keeping the transcription's timings, which fixes brand-name spellings.
If the clip has no speech, a speech job fails with caption_no_speech; send cues with text, start and end instead.
| Field | What it does | Use it when |
|---|---|---|
| style | Look and motion; default slam for Latin | You want a different look |
| design.phrasing.max_words | Words per caption card | Small phone screens |
| script_text | Align burned text to your script | Brand names are misheard |
| cues | Authored text with start and end, no speech to text | The clip is silent |
| language | Speech-to-text hint | The language is not auto-detected |
The request
The input must be a public HTTPS video URL. The default mode is async, so keep the returned job id and read the result when it is ready.
import json, os, urllib.request
body = {
"video_url": "https://media.sume.com/artifacts/artf_demo/ad-15s.mp4",
"style": "slam",
"design": {"phrasing": {"max_words": 3}},
}
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
data=json.dumps(body).encode(),
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": "reddit15-captions-001",
},
)
with urllib.request.urlopen(req) as r:
print(json.load(r))
Does this work for languages other than English?
Yes, with one rule. Korean speech needs a Hangul style. If you send Korean text to slam, punch or tiktok-green, the API answers 400 with caption_hangul_text_latin_style instead of silently rendering boxes. Use black-outline, the safe default, or korean-ad for the ad karaoke look. If you leave style unset, the text decides: Korean resolves to black-outline.
language is only a hint for speech-to-text. It never picks a style or a font, so a Korean ad still needs a Korean-capable style chosen on purpose. For Korean you can also pass font, for example pretendard for a clean baseline or do-hyeon for the thick CapCut feel, and Sume rejects a Hangul font on a Latin style.
One more setting is worth knowing for a feed ad. design.placement.anchor_ratio sets where the caption line sits as a fraction of the frame height, and landscape_anchor_ratio does the same for wide frames. Moving the line away from where a platform overlays its own buttons is a design choice you make once and reuse across the 15-second cuts.
A check worth doing
Read the first caption on a phone-sized preview. If your hook is in the first second, the first card must appear in that second. If it does not, the speech starts late in the file and the trim's start should move, not the caption.
Sources
Related posts
More in Use cases
- Reddit 15-second ad: find the best 15 seconds with video frames
Before cutting a long ad to 15 seconds for Reddit, pull up to 24 stills with POST /v1/video-frames, one every two seconds, and choose where the cut starts.
- Reddit's 15-second video views: what a 15 s clip costs on Sume models
Reddit bills videos over 15 s at 15 s in its Engaged Video Views beta. A 15-second clip costs $1.88 on Wan 3.0 at 720p or $2.10 on Kling 3 without sound.
- Trim 50 ads to 15 seconds for Reddit: the batch costs $1.00
Fifty finished ads cut to 15 seconds with /v1/video-trim at $0.02 a job is $1.00 before retries. A Python loop with one idempotency key per ad.
- Restaurant table tent art: Ideogram 4.5 at 3:4, specials quoted
Make table tent artwork for a restaurant: a 3:4 Ideogram 4.5 image with each special quoted in the prompt, a price check, and the real cost per tent on Sume.
Written by Sume