Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.

If burned captions disappear against bright or busy footage, darken the clip first with a dim op on Video filter, then burn captions onto the dimmed MP4 and adjust colours with design. Two jobs, about $0.02 plus $0.20 under the current published rates.
Why dim before captioning rather than after?
The caption renderer draws text onto whatever pixels are underneath. Dimming multiplies luma across the whole clip, so the text keeps its colour while the footage behind it gets darker. Black stays black and chroma does not shift. The amount is in the range (0, 1]: 0.45 is noticeably darker and 1 is unchanged. 0 and anything above 1 are refused.
Dim is a whole-clip pass. Sume does not darken only the area behind the text, so the footage looks moodier overall. That suits faceless B-roll and does not suit product shots where colour accuracy matters.
What is the two-step call?
Import the clip first (POST /v1/media-imports), because Video filter only reads this workspace's media.sume.com media. The program can be checked for free with POST /v1/video-filter/check before you pay for the encode.
curl -X POST https://api.sume.com/v1/video-filter \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dim-broll-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/broll.mp4",
"ops": [{ "op": "dim", "amount": 0.45 }]
}'When the job is ready, GET /v1/jobs/:id/result returns the dimmed video_url. Pass that URL to the caption call, with a design override to set the colours:
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-dimmed-001" \
-d '{
"video_url": "<video_url from the dim result>",
"style": "black-outline",
"design": { "colors": { "base": "#FFFFFF", "stroke": "#000000" } }
}'Which caption knobs help legibility?
Each field is optional and merges over the style's own value. Numbers outside a documented range return a 400 at request time, so a bad value fails before it bills; check the ranges in the OpenAPI schema before pushing weights or stroke widths. Colours are hex, rgb()/rgba() or transparent.
design is not supported on punch or tiktok-green, which render on a path that reads none of these tokens. Use slam for Latin text or black-outline for Korean.
| Group | Fields | Use |
|---|---|---|
| colors | base, active, stroke, accent, card | Light text, dark outline, optional card behind the line |
| typography | base_weight, font_size_ratio, stroke_width_px | Heavier and larger text |
| placement | anchor_ratio, landscape_anchor_ratio | Move the line to a calmer part of the frame |
| phrasing | max_words, max_chars, pause_seconds | Shorter lines |
How do you choose the dim amount?
Start at 0.45 and judge on a phone. Light footage such as snow, beaches or white backgrounds usually needs a lower value, and already dark footage needs none. Because the range is (0, 1], you can step through 0.7, 0.55 and 0.45 and compare. Each encode is a billed job, so run POST /v1/video-filter/check first: it validates the program for free and returns diagnostics instead of a 400.
A passing check does not promise the encode succeeds, since a bad expression or a memory limit can still fail on the worker, but those come back as a structured job error.
Is there a cheaper path?
If only a few clips are hard to read, dim only those. If every clip in the video is busy, consider a calmer placement instead: placement.anchor_ratio moves the line centre as a fraction of frame height, and landscape_anchor_ratio does the same for wide frames. Moving the text to a quiet part of the picture costs nothing extra beyond the $0.20 caption job, and it keeps the colours of the footage intact. Whichever route you take, keep the caption job's idempotency key distinct per attempt so a retry does not return the old render.
What does this not guarantee?
Sume does not measure contrast ratios, and nothing here is a statement about WCAG compliance. Look at the result on a phone, in sunlight if you can, before publishing. The dim step also has a limit: the source must be 300 seconds or shorter, so long videos need a trim first.
Sources
Related posts
More in Media tools
- Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.
- Check an AI take says your script: hypit align and unmatched words
hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.
- Join voiceover takes into one gapless track with Timeline audio
Concatenate up to 20 Sume-hosted voice takes with Timeline audio, with no seam silence and no re-synthesis, and re-base video starts from the returned offsets.
- Crop a 16:9 video to a centered 9:16 strip: the crop fractions
For a centered 9:16 crop of a 16:9 video, send crop x 0.3418, y 0, width 0.3164, height 1 to Sume video-filter. 1:1 and 4:5 values are in the table.
Written by Sume