Caption a talking-head video: style, placement and phrasing
Caption a talking-head clip with the Sume API: pick a style, move the line off the face, set words per card. One request, $0.20 for up to 60 seconds.

For a talking-head video, send one POST /v1/video-captions request with the clip's public HTTPS video_url, pick a style, and use design to move the line away from the face and shorten each phrase card. A standalone caption job is $0.20 for videos of up to 60 seconds, and the speech-to-text step is included, so you do not transcribe separately unless you want the text.
The defaults are reasonable. If you omit style, Latin text resolves to slam and Korean to black-outline. The rest of this page is about the three knobs that matter most when a person's face fills the frame: look, position and pacing.
Pick the look
The documented Latin styles are slam, punch and tiktok-green. The Hangul styles are korean-ad, black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis. Two rules save you money: punch and tiktok-green render on a path that reads no design tokens, so sending design with them is refused, and Korean text on a Latin style returns 400 caption_hangul_text_latin_style before any render.
The language field is only a hint to speech-to-text. It never picks the style or the font, so a Korean talker needs a Hangul style chosen explicitly, or the Korean default applies when the resolved text is Korean.
Move the line off the face
design.placement.anchor_ratio is the center of the caption line as a fraction of the frame height, and the documented range is 0.05 to 0.95. On a vertical talking head the face usually sits in the upper half, so a value around 0.7 to 0.75 puts the line lower in the frame, below the mouth. landscape_anchor_ratio does the same for 16:9 clips, which tend to frame a person lower in the picture.
Treat those numbers as a starting point and look at one render. Sume documents the range, not a face-safe value.
| Field | Range | What it changes |
|---|---|---|
| placement.anchor_ratio | 0.05 to 0.95 | Vertical center of the caption line, as a fraction of frame height |
| placement.landscape_anchor_ratio | 0.05 to 0.95 | Same, for landscape output |
| phrasing.max_words | 1 to 12 | Words per phrase card |
| phrasing.max_chars | 4 to 60 | Characters per card |
| phrasing.pause_seconds | 0.05 to 3 | Silence that ends a card |
| typography.safe_width_ratio | 0.3 to 1 | How wide the text may run |
Set the pacing
Fast talkers need short cards. Lower max_words to 3 or 4 and the line changes more often, with fewer words to read at once; raise it for a calm explainer. pause_seconds decides how long a silence must last before a card ends, so a speaker who pauses between thoughts gets clean breaks if you keep it small.
None of these values is measured against viewer behavior in Sume's docs. Use them to match the speaker's rhythm, then verify by watching the result.
The request
This uses slam, a lower anchor and short cards. Add language only when auto-detect misses.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: talking-head-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/head.mp4",
"style": "slam",
"design": {
"placement": { "anchor_ratio": 0.72 },
"phrasing": { "max_words": 3, "pause_seconds": 0.3 }
}
}'Fixing it without paying twice for speech
If the first render puts the line in the wrong place, do not resubmit the URL. Send source_caption_id with the new design. Sume reuses the source video and the word timings it already has, so speech-to-text does not run again, though a restyle is still a render at the same price. To fix a misheard word, send words with the correction on the same source_caption_id call.
Poll GET /v1/jobs/:id/status until completed, then read the captioned video_url from GET /v1/video-captions/:id.
Three mistakes to avoid
The first is sending design with punch or tiktok-green. Those two styles ignore design tokens, and the API refuses the combination at request time, so you waste a call instead of a render. If you want an accent color change, pick slam or a Hangul style and set design.colors.active.
The second is trusting language to choose a font. It does not. A Korean talker on a Latin style is rejected with caption_hangul_text_latin_style, and a Hangul font on a Latin style is rejected with caption_font_requires_hangul_style.
The third is captioning a clip with no speech. If the person is silent for the whole clip, speech-to-text finds nothing and the job fails with caption_no_speech; send cues for text you wrote yourself.
Sources
Related posts
More in Media tools
- Captions for a video with loud background music: script text or cues
Music can bury speech and trip speech-to-text. Sume documents three ways to supply your own wording: script_text, a words array or cues. Which to pick.
- Crop 16:9 to 9:16 with video filter: the fractions to send
For a 16:9 source, a centred 9:16 crop is x 0.3418, y 0, width 0.3164, height 1. One $0.02 video-filter job, checked free first. Examples for 1920x1080.
- How many 6-second clips fill a 3-minute Short, Reel or TikTok ad?
A 3-minute Short needs 30 six-second clips, a 20-minute Reel 200, a 10-minute TikTok ad 100. Slot limits, chunked renders and Timeline cost for each.
- What prompt makes a 15-second Black Friday ad track?
Write the track as three timed sections: build, drop, sting. The Music Router takes it for a flat $0.125; the whole 15-second ad is $2.305 on Sume.
Written by Sume