Avatar video captions: inline has four knobs, no design override

Inline captions on a Sume avatar video take style, font, language and script_text. For design colors, caption the clean MP4 with standalone captions.

5 min readSume
All posts

Inline captions on a Sume avatar video accept the same four knobs as standalone video captions: style, an optional font, a language hint and script_text. The per-request design overrides for colours, typography, placement, phrasing and motion are documented on the standalone caption job, so to tune them render the avatar video clean and caption it afterwards.

That is two jobs instead of one, but it keeps the clean MP4 as your master.

What does inline captioning do?

Set captions on POST /v1/avatar-1.0/talking-video and Sume burns the style into the final MP4 after generation, using the spoken script or video_inputs text. Styles are slam (the default), punch, tiktok-green and korean-ad, plus the Hangul identities weight-shift, black-outline, highlight, pill-karaoke, clip-wipe and editorial-emphasis.

Inline captions do not create a separate billed video-caption job. If the estimated duration is above 60 seconds they are rejected, and a caption-stage failure soft-fails: the avatar job can still succeed with a clean video_url and captions.status=failed.

{
  "avatar_handle": "sume_clawra",
  "script": "Three things to check before you launch.",
  "captions": { "enabled": true, "style": "slam", "language": "auto" }
}

What can only the standalone job change?

The standalone caption docs define design: colours (base, active, stroke, accent, accent_deep, card), typography, placement anchors, phrasing (max_words, max_chars, pause_seconds) and motion timings. Every field is optional and merges over the chosen style. Numbers outside their documented range return a 400 instead of rendering wrong. design is not supported on punch or tiktok-green.

Inline versus standalone captions (Sume docs, read 2026-10-02)
QuestionInline on avatar videoStandalone video captions
InputScript or video_inputs of the avatar jobPublic HTTPS video_url
Knobs documentedstyle, font, language, script_textSame, plus design, words, cues and segments
BillingNo separate caption jobIts own job
On failureSoft fail, clean video keptA normal job result

How do I caption a finished avatar video with a custom colour?

Fetch the completed job result, take the public media.sume.com video URL, and send it to POST /v1/video-captions with a style and a design override. The docs' own example keeps black-outline and changes only the active colour to cyan.

Keep the clean render. If the caption look needs another pass you re-burn from the clean file rather than rendering the avatar again.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recaption-avatar-001" \
  -d '{"video_url":"https://media.sume.com/artifacts/your-clean-video.mp4","style":"black-outline","design":{"colors":{"active":"#22D3EE"}}}'

Which style should I pick for Korean speech?

Not slam, punch or tiktok-green. Their Latin display faces have no Hangul glyphs, so Korean copy returns 400 caption_hangul_text_latin_style rather than being re-styled, and naming a Hangul font with a Latin style returns 400 caption_font_requires_hangul_style. Use a Hangul identity such as black-outline or korean-ad and pair it with language: "ko".

For English scripts, slam is the default and a safe starting point.

Which path should I choose?

Use inline when the stock look is fine and you want one job. Use two steps when brand colours or phrasing matter, or when you may re-caption for another language without regenerating the avatar. Preview stills are never captioned in either case; captions stored on a preview apply only at generate-video. For failure handling, see avatar video captions fail but the video succeeds.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume