Caption a talking-head video: style, placement and phrasing

Caption a talking-head clip with the Sume API: pick a style, move the line off the face, set words per card. One request, $0.20 for up to 60 seconds.

5 min readSume
All posts

For a talking-head video, send one POST /v1/video-captions request with the clip's public HTTPS video_url, pick a style, and use design to move the line away from the face and shorten each phrase card. A standalone caption job is $0.20 for videos of up to 60 seconds, and the speech-to-text step is included, so you do not transcribe separately unless you want the text.

The defaults are reasonable. If you omit style, Latin text resolves to slam and Korean to black-outline. The rest of this page is about the three knobs that matter most when a person's face fills the frame: look, position and pacing.

Pick the look

The documented Latin styles are slam, punch and tiktok-green. The Hangul styles are korean-ad, black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis. Two rules save you money: punch and tiktok-green render on a path that reads no design tokens, so sending design with them is refused, and Korean text on a Latin style returns 400 caption_hangul_text_latin_style before any render.

The language field is only a hint to speech-to-text. It never picks the style or the font, so a Korean talker needs a Hangul style chosen explicitly, or the Korean default applies when the resolved text is Korean.

Move the line off the face

design.placement.anchor_ratio is the center of the caption line as a fraction of the frame height, and the documented range is 0.05 to 0.95. On a vertical talking head the face usually sits in the upper half, so a value around 0.7 to 0.75 puts the line lower in the frame, below the mouth. landscape_anchor_ratio does the same for 16:9 clips, which tend to frame a person lower in the picture.

Treat those numbers as a starting point and look at one render. Sume documents the range, not a face-safe value.

Design fields that matter for a talking head, ranges from the Sume caption request schema read 2026-10-07
FieldRangeWhat it changes
placement.anchor_ratio0.05 to 0.95Vertical center of the caption line, as a fraction of frame height
placement.landscape_anchor_ratio0.05 to 0.95Same, for landscape output
phrasing.max_words1 to 12Words per phrase card
phrasing.max_chars4 to 60Characters per card
phrasing.pause_seconds0.05 to 3Silence that ends a card
typography.safe_width_ratio0.3 to 1How wide the text may run

Set the pacing

Fast talkers need short cards. Lower max_words to 3 or 4 and the line changes more often, with fewer words to read at once; raise it for a calm explainer. pause_seconds decides how long a silence must last before a card ends, so a speaker who pauses between thoughts gets clean breaks if you keep it small.

None of these values is measured against viewer behavior in Sume's docs. Use them to match the speaker's rhythm, then verify by watching the result.

The request

This uses slam, a lower anchor and short cards. Add language only when auto-detect misses.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: talking-head-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/head.mp4",
    "style": "slam",
    "design": {
      "placement": { "anchor_ratio": 0.72 },
      "phrasing": { "max_words": 3, "pause_seconds": 0.3 }
    }
  }'

Fixing it without paying twice for speech

If the first render puts the line in the wrong place, do not resubmit the URL. Send source_caption_id with the new design. Sume reuses the source video and the word timings it already has, so speech-to-text does not run again, though a restyle is still a render at the same price. To fix a misheard word, send words with the correction on the same source_caption_id call.

Poll GET /v1/jobs/:id/status until completed, then read the captioned video_url from GET /v1/video-captions/:id.

Three mistakes to avoid

The first is sending design with punch or tiktok-green. Those two styles ignore design tokens, and the API refuses the combination at request time, so you waste a call instead of a render. If you want an accent color change, pick slam or a Hangul style and set design.colors.active.

The second is trusting language to choose a font. It does not. A Korean talker on a Latin style is rejected with caption_hangul_text_latin_style, and a Hangul font on a Latin style is rejected with caption_font_requires_hangul_style.

The third is captioning a clip with no speech. If the person is silent for the whole clip, speech-to-text finds nothing and the job fails with caption_no_speech; send cues for text you wrote yourself.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume