Avatar video captions: inline has four knobs, no design override
Inline captions on a Sume avatar video take style, font, language and script_text. For design colors, caption the clean MP4 with standalone captions.
Inline captions on a Sume avatar video accept the same four knobs as standalone video captions: style, an optional font, a language hint and script_text. The per-request design overrides for colours, typography, placement, phrasing and motion are documented on the standalone caption job, so to tune them render the avatar video clean and caption it afterwards.
That is two jobs instead of one, but it keeps the clean MP4 as your master.
What does inline captioning do?
Set captions on POST /v1/avatar-1.0/talking-video and Sume burns the style into the final MP4 after generation, using the spoken script or video_inputs text. Styles are slam (the default), punch, tiktok-green and korean-ad, plus the Hangul identities weight-shift, black-outline, highlight, pill-karaoke, clip-wipe and editorial-emphasis.
Inline captions do not create a separate billed video-caption job. If the estimated duration is above 60 seconds they are rejected, and a caption-stage failure soft-fails: the avatar job can still succeed with a clean video_url and captions.status=failed.
{
"avatar_handle": "sume_clawra",
"script": "Three things to check before you launch.",
"captions": { "enabled": true, "style": "slam", "language": "auto" }
}What can only the standalone job change?
The standalone caption docs define design: colours (base, active, stroke, accent, accent_deep, card), typography, placement anchors, phrasing (max_words, max_chars, pause_seconds) and motion timings. Every field is optional and merges over the chosen style. Numbers outside their documented range return a 400 instead of rendering wrong. design is not supported on punch or tiktok-green.
| Question | Inline on avatar video | Standalone video captions |
|---|---|---|
| Input | Script or video_inputs of the avatar job | Public HTTPS video_url |
| Knobs documented | style, font, language, script_text | Same, plus design, words, cues and segments |
| Billing | No separate caption job | Its own job |
| On failure | Soft fail, clean video kept | A normal job result |
How do I caption a finished avatar video with a custom colour?
Fetch the completed job result, take the public media.sume.com video URL, and send it to POST /v1/video-captions with a style and a design override. The docs' own example keeps black-outline and changes only the active colour to cyan.
Keep the clean render. If the caption look needs another pass you re-burn from the clean file rather than rendering the avatar again.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recaption-avatar-001" \
-d '{"video_url":"https://media.sume.com/artifacts/your-clean-video.mp4","style":"black-outline","design":{"colors":{"active":"#22D3EE"}}}'Which style should I pick for Korean speech?
Not slam, punch or tiktok-green. Their Latin display faces have no Hangul glyphs, so Korean copy returns 400 caption_hangul_text_latin_style rather than being re-styled, and naming a Hangul font with a Latin style returns 400 caption_font_requires_hangul_style. Use a Hangul identity such as black-outline or korean-ad and pair it with language: "ko".
For English scripts, slam is the default and a safe starting point.
Which path should I choose?
Use inline when the stock look is fine and you want one job. Use two steps when brand colours or phrasing matter, or when you may re-caption for another language without regenerating the avatar. Preview stills are never captioned in either case; captions stored on a preview apply only at generate-video. For failure handling, see avatar video captions fail but the video succeeds.
Sources
Related posts
More in Sume Avatar 1.0
- Sume avatar video stuck? Read the job events before retrying
Use GET /v1/jobs/{id}/events to see whether an avatar video is queued, started, submitted or failed before you cancel, wait, or resubmit and pay twice.
- Avatar video soundtrack: send prompt or audio_url, volume 0.05-0.4
Sume's Avatar Video package accepts a soundtrack with exactly one of prompt or audio_url and a volume from 0.05 to 0.4, default 0.15. What each costs and does.
- Avatar video preview: approve the first frame, then pick quality
Sume avatar video previews are tier-independent: approve stills, then render at standard, plus or max. What can change at generate-video, and what cannot.
- Avatar preview resource_status vs job_status: which field to poll
Avatar video previews return resource_status and job_status next to a legacy status field. Read resource_status for readiness and job_status for polling.
Written by Sume