Reddit 15-second video ad: burn captions so it reads on mute

For a 15-second Reddit video ad, burn captions into the cut with POST /v1/video-captions: one $0.20 job, style slam for English, no new edit.

5 min readSume
All posts

To make a 15-second Reddit video ad readable with the sound off, run the trimmed cut through POST /v1/video-captions, which burns captions into the picture and returns a new video. For English speech, leave style unset and Sume picks slam; the job is priced at $0.20 for clips up to 60 seconds. Reddit's 15-second Engaged Video Views option is listed as a public beta from 2026-09-01, with longer videos billing at 15 s (read 2026-10-06), so the captions only need to cover those 15 seconds.

Facts: Video captions: burn captions onto a video for the endpoint and price, and the Social media platform updates, October 2026 (read 2026-10-06) for the Reddit bid.

Why caption after the trim, not before?

Caption after you trim. A caption job transcribes the clip it receives, so a 15-second input is a 15-second job and nothing is burned for seconds that will not be bought. If you caption the 30-second master first and trim after, you pay the same $0.20 but you must also keep the caption timing aligned to the new start time. Order the steps trim, then captions.

Which options matter for a short ad?

Three fields do most of the work. style sets the look and motion; unset, Latin text resolves to slam and Korean text to black-outline. design overrides one request's colors, typography, placement and phrasing, for example phrasing.max_words to keep two or three words on screen. script_text makes the burned text match your script while keeping the transcription's timings, which fixes brand-name spellings.

If the clip has no speech, a speech job fails with caption_no_speech; send cues with text, start and end instead.

Caption fields for a short ad, from Sume's video captions page, read 2026-10-06.
FieldWhat it doesUse it when
styleLook and motion; default slam for LatinYou want a different look
design.phrasing.max_wordsWords per caption cardSmall phone screens
script_textAlign burned text to your scriptBrand names are misheard
cuesAuthored text with start and end, no speech to textThe clip is silent
languageSpeech-to-text hintThe language is not auto-detected

The request

The input must be a public HTTPS video URL. The default mode is async, so keep the returned job id and read the result when it is ready.

import json, os, urllib.request

body = {
    "video_url": "https://media.sume.com/artifacts/artf_demo/ad-15s.mp4",
    "style": "slam",
    "design": {"phrasing": {"max_words": 3}},
}
req = urllib.request.Request(
    "https://api.sume.com/v1/video-captions",
    data=json.dumps(body).encode(),
    headers={
        "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Content-Type": "application/json",
        "Idempotency-Key": "reddit15-captions-001",
    },
)
with urllib.request.urlopen(req) as r:
    print(json.load(r))

Does this work for languages other than English?

Yes, with one rule. Korean speech needs a Hangul style. If you send Korean text to slam, punch or tiktok-green, the API answers 400 with caption_hangul_text_latin_style instead of silently rendering boxes. Use black-outline, the safe default, or korean-ad for the ad karaoke look. If you leave style unset, the text decides: Korean resolves to black-outline.

language is only a hint for speech-to-text. It never picks a style or a font, so a Korean ad still needs a Korean-capable style chosen on purpose. For Korean you can also pass font, for example pretendard for a clean baseline or do-hyeon for the thick CapCut feel, and Sume rejects a Hangul font on a Latin style.

One more setting is worth knowing for a feed ad. design.placement.anchor_ratio sets where the caption line sits as a fraction of the frame height, and landscape_anchor_ratio does the same for wide frames. Moving the line away from where a platform overlays its own buttons is a design choice you make once and reuse across the 15-second cuts.

A check worth doing

Read the first caption on a phone-sized preview. If your hook is in the first second, the first card must appear in that second. If it does not, the speech starts late in the file and the trim's start should move, not the caption.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume