Avatar video captions: inline add-on or standalone $0.20 job?

Inline avatar-video captions create no separate caption job; standalone captions reserve $0.20 per video up to 60 seconds. Compare limits and styles.

5 min readSume
All posts

Inline captions on an Avatar Video are an add-on on the avatar-video estimate and create no separate caption job. A standalone caption job reserves and captures $0.20 of Sume usage for videos up to 60 seconds under the current fixed estimate.

Both facts are on the Video captions page, read 2026-10-02, which also says to confirm live pricing in GET /v1/catalog and the OpenAPI. The docs do not publish the inline add-on's dollar amount, so this post does not quote one.

What is the difference in one table?

Both paths use the same styles, so the choice is mostly about when you decide to caption.

Inline versus standalone captions (read 2026-10-02)
Inline on avatar videoStandalone video captions
Requestcaptions block on talking-video or previewPOST /v1/video-captions with video_url
BillingAdd-on on the avatar-video estimate$0.20 per video up to 60 seconds
ResourceNo video_caption resourcevideo_caption resource
DurationEstimate above 60 seconds is rejectedFixed estimate stated for up to 60 seconds
Design overridesNot availabledesign field available
On failureSoft-fail: clean video kept, captions.status failedNormal job failure

When should I turn on inline captions?

Choose inline when the video is new, the script is the source of the captions and the defaults suit you. The captions burn into the clean final MP4 after generation from the spoken script or video_inputs text, so there is no second upload and no second job to track.

Inline also fits a preview-first workflow: caption intent stored on preview create applies at generate-video. The stills are never captioned.

{
  "avatar_handle": "sume_clawra",
  "script": "Three tips for a calmer morning.",
  "captions": {"enabled": true, "style": "slam", "language": "auto"}
}

When is the standalone job the better buy?

Choose standalone when you already have a finished clip, need design overrides such as a different highlight color, or want to restyle after the fact. It takes a public HTTPS video_url, and a silent clip fails as caption_no_speech with next_action: use_overlay_captions; pass cues with text, start and end to burn authored copy without speech-to-text.

If inline captions fail, you still have the clean video, and a standalone job on its public URL is the recovery. That path costs the fixed $0.20 for up to 60 seconds, so budget for it if your batch is large.

Korean and other limits to remember

Latin-only styles are rejected for Korean script on the inline path with a 400, so name a Hangul style there. Standalone captions resolve an omitted style by wording: slam for Latin, black-outline for Korean.

Neither path accepts SRT uploads; phrase-level text goes in cues or segments. For the tradeoff in dollars, read your own GET /v1/usage after a test render rather than assuming.

How do I decide in practice?

Ask three questions. Is the video still to be rendered? Then inline is the simplest path. Do you need design overrides or a restyle? Then standalone. Is the clip longer than 60 seconds by estimate? Then inline is off the table, and the standalone price is stated only for videos up to 60 seconds.

If you are unsure, render without captions first and caption the clean MP4 afterwards. You pay the fixed standalone amount, but you keep full control of the look, and the clean video is never at risk.

Quick answers

Do previews show captions? No: preview stills are never captioned, and caption intent stored on a preview applies only when you call generate-video. Does an inline caption failure fail the job? No: the avatar job can still succeed with a clean video_url and captions.status=failed, so check that field separately. Where do I read what I actually spent? GET /v1/usage and your dashboard are the billing records, not an estimate in a blog post.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume