Move burned-in captions up or down: restyle with source_caption_id
Change caption placement with design.placement.anchor_ratio and restyle with source_caption_id: the same word timings, no second speech-to-text run.

To move burned-in captions higher or lower on an existing captioned clip, send a new video-captions request with source_caption_id instead of video_url and set design.placement.anchor_ratio, the centre of the caption line as a fraction of the frame height. Sume reuses the source video of that caption and the word timings it already has, so speech-to-text does not run a second time. The docs say the price does not change, because a restyle is still a render.
That makes placement a cheap thing to iterate on: you can test a few positions against a platform's safe-zone file without paying for a new transcription each time.
The design fields
A style is a set of design tokens and design overrides them for one request. Sume merges each field over the style's value, so one key changes one thing. The groups are colors, typography, placement, phrasing and motion. Under placement there are two: anchor_ratio and landscape_anchor_ratio.
Typography has font_size_ratio and safe_width_ratio, which shape how large the text is and how wide a line may run. A number outside its documented range returns a 400, so a bad value fails at request time and you do not pay for an incorrect render. Note that punch and tiktok-green do not support design.
| Group | Fields |
|---|---|
| placement | anchor_ratio, landscape_anchor_ratio |
| typography | base_weight, active_weight, active_scale, font_size_ratio, safe_width_ratio, stroke_width_px |
| phrasing | max_words, max_chars, pause_seconds |
| colors | base, active, stroke, accent, accent_deep, card |
| motion | enter_seconds, exit_seconds, emphasis_in_seconds, emphasis_out_seconds |
A restyle request
The request below restyles a finished caption job with the black-outline style and sets anchor_ratio to 0.62. The docs define anchor_ratio only as the centre of the line as a fraction of the frame height, so confirm where 0.62 lands by pulling a frame from the result.
The new job is billed as a caption job: $0.20 per job for a clip of up to 60 seconds.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-restyle-001" \
-d '{
"source_caption_id": "cap_demo",
"style": "black-outline",
"design": {
"placement": { "anchor_ratio": 0.62 },
"typography": { "safe_width_ratio": 0.8 }
}
}'Checking where it landed
After the job completes, call video frames or video inspect with an at time inside a caption and look at the still. The related post on spot-checking burned-in captions shows the recipe. Do this for each placement you try; three placements cost three caption jobs, 3 x $0.20 = $0.60 for a clip of up to 60 seconds.
If the captions must clear the on-screen interface of a platform, use that platform's own safe-zone file as the reference, and check the frames against it rather than relying on a guessed ratio.
Corrections without re-transcribing
If the spoken words were transcribed wrongly, send words with the restyle to correct the text. Sending only one of script_text, words, cues and segments is allowed per request. With words or cues, Sume does not run speech-to-text at all and burns your text at the times you give.
A practical order of work
Caption the clip once with the default style and keep the job id: that is the one speech-to-text run you pay for, $0.20 for a clip of 60 seconds or less. Then restyle with source_caption_id for each look you want to compare and keep the best. Each restyle is a render at the same job price, so four looks cost 5 x $0.20 = $1.00 in all, counting the first.
Keep the design overrides in a small table in your own code, one row per look, so the same overrides can be applied to every clip in a batch. Because design merges over the style, a row can be as short as one placement value.
Sources
Related posts
More in Media tools
- One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
Written by Sume