Captions unreadable on busy footage: dim the clip, then burn

Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.

4 min readSume
All posts

If burned captions disappear against bright or busy footage, darken the clip first with a dim op on Video filter, then burn captions onto the dimmed MP4 and adjust colours with design. Two jobs, about $0.02 plus $0.20 under the current published rates.

Why dim before captioning rather than after?

The caption renderer draws text onto whatever pixels are underneath. Dimming multiplies luma across the whole clip, so the text keeps its colour while the footage behind it gets darker. Black stays black and chroma does not shift. The amount is in the range (0, 1]: 0.45 is noticeably darker and 1 is unchanged. 0 and anything above 1 are refused.

Dim is a whole-clip pass. Sume does not darken only the area behind the text, so the footage looks moodier overall. That suits faceless B-roll and does not suit product shots where colour accuracy matters.

What is the two-step call?

Import the clip first (POST /v1/media-imports), because Video filter only reads this workspace's media.sume.com media. The program can be checked for free with POST /v1/video-filter/check before you pay for the encode.

curl -X POST https://api.sume.com/v1/video-filter \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dim-broll-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/broll.mp4",
    "ops": [{ "op": "dim", "amount": 0.45 }]
  }'

When the job is ready, GET /v1/jobs/:id/result returns the dimmed video_url. Pass that URL to the caption call, with a design override to set the colours:

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-dimmed-001" \
  -d '{
    "video_url": "<video_url from the dim result>",
    "style": "black-outline",
    "design": { "colors": { "base": "#FFFFFF", "stroke": "#000000" } }
  }'

Which caption knobs help legibility?

Each field is optional and merges over the style's own value. Numbers outside a documented range return a 400 at request time, so a bad value fails before it bills; check the ranges in the OpenAPI schema before pushing weights or stroke widths. Colours are hex, rgb()/rgba() or transparent.

design is not supported on punch or tiktok-green, which render on a path that reads none of these tokens. Use slam for Latin text or black-outline for Korean.

Design groups relevant to legibility (read 2026-10-02)
GroupFieldsUse
colorsbase, active, stroke, accent, cardLight text, dark outline, optional card behind the line
typographybase_weight, font_size_ratio, stroke_width_pxHeavier and larger text
placementanchor_ratio, landscape_anchor_ratioMove the line to a calmer part of the frame
phrasingmax_words, max_chars, pause_secondsShorter lines

How do you choose the dim amount?

Start at 0.45 and judge on a phone. Light footage such as snow, beaches or white backgrounds usually needs a lower value, and already dark footage needs none. Because the range is (0, 1], you can step through 0.7, 0.55 and 0.45 and compare. Each encode is a billed job, so run POST /v1/video-filter/check first: it validates the program for free and returns diagnostics instead of a 400.

A passing check does not promise the encode succeeds, since a bad expression or a memory limit can still fail on the worker, but those come back as a structured job error.

Is there a cheaper path?

If only a few clips are hard to read, dim only those. If every clip in the video is busy, consider a calmer placement instead: placement.anchor_ratio moves the line centre as a fraction of frame height, and landscape_anchor_ratio does the same for wide frames. Moving the text to a quiet part of the picture costs nothing extra beyond the $0.20 caption job, and it keeps the colours of the footage intact. Whichever route you take, keep the caption job's idempotency key distinct per attempt so a retry does not return the old render.

What does this not guarantee?

Sume does not measure contrast ratios, and nothing here is a statement about WCAG compliance. Look at the result on a phone, in sunlight if you can, before publishing. The dim step also has a limit: the source must be 300 seconds or shorter, so long videos need a trim first.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume