Avatar video captions: inline add-on or standalone $0.20 job?
Inline avatar-video captions create no separate caption job; standalone captions reserve $0.20 per video up to 60 seconds. Compare limits and styles.
Inline captions on an Avatar Video are an add-on on the avatar-video estimate and create no separate caption job. A standalone caption job reserves and captures $0.20 of Sume usage for videos up to 60 seconds under the current fixed estimate.
Both facts are on the Video captions page, read 2026-10-02, which also says to confirm live pricing in GET /v1/catalog and the OpenAPI. The docs do not publish the inline add-on's dollar amount, so this post does not quote one.
What is the difference in one table?
Both paths use the same styles, so the choice is mostly about when you decide to caption.
| Inline on avatar video | Standalone video captions | |
|---|---|---|
| Request | captions block on talking-video or preview | POST /v1/video-captions with video_url |
| Billing | Add-on on the avatar-video estimate | $0.20 per video up to 60 seconds |
| Resource | No video_caption resource | video_caption resource |
| Duration | Estimate above 60 seconds is rejected | Fixed estimate stated for up to 60 seconds |
| Design overrides | Not available | design field available |
| On failure | Soft-fail: clean video kept, captions.status failed | Normal job failure |
When should I turn on inline captions?
Choose inline when the video is new, the script is the source of the captions and the defaults suit you. The captions burn into the clean final MP4 after generation from the spoken script or video_inputs text, so there is no second upload and no second job to track.
Inline also fits a preview-first workflow: caption intent stored on preview create applies at generate-video. The stills are never captioned.
{
"avatar_handle": "sume_clawra",
"script": "Three tips for a calmer morning.",
"captions": {"enabled": true, "style": "slam", "language": "auto"}
}When is the standalone job the better buy?
Choose standalone when you already have a finished clip, need design overrides such as a different highlight color, or want to restyle after the fact. It takes a public HTTPS video_url, and a silent clip fails as caption_no_speech with next_action: use_overlay_captions; pass cues with text, start and end to burn authored copy without speech-to-text.
If inline captions fail, you still have the clean video, and a standalone job on its public URL is the recovery. That path costs the fixed $0.20 for up to 60 seconds, so budget for it if your batch is large.
Korean and other limits to remember
Latin-only styles are rejected for Korean script on the inline path with a 400, so name a Hangul style there. Standalone captions resolve an omitted style by wording: slam for Latin, black-outline for Korean.
Neither path accepts SRT uploads; phrase-level text goes in cues or segments. For the tradeoff in dollars, read your own GET /v1/usage after a test render rather than assuming.
How do I decide in practice?
Ask three questions. Is the video still to be rendered? Then inline is the simplest path. Do you need design overrides or a restyle? Then standalone. Is the clip longer than 60 seconds by estimate? Then inline is off the table, and the standalone price is stated only for videos up to 60 seconds.
If you are unsure, render without captions first and caption the clean MP4 afterwards. You pay the fixed standalone amount, but you keep full control of the look, and the clean video is never at risk.
Quick answers
Do previews show captions? No: preview stills are never captioned, and caption intent stored on a preview applies only when you call generate-video. Does an inline caption failure fail the job? No: the avatar job can still succeed with a clean video_url and captions.status=failed, so check that field separately. Where do I read what I actually spent? GET /v1/usage and your dashboard are the billing records, not an estimate in a blog post.
Sources
Related posts
More in Pricing
- Does a product image cost extra in a Sume avatar video?
Yes: product_image raises the per-second rate by $0.010 on standard, $0.013 on plus and $0.030 on max. The numbers, a 30-second example, and when to skip it.
- Azure Speech free tier: 5 audio hours is $3.00 on Sume STT
Azure's F0 tier gives 5 audio hours of speech-to-text and 0.5M TTS characters a month. At Sume's $0.01 a minute, 5 hours costs $3.00.
- Black Friday price variants: one clip, twelve caption jobs, $2.40
Test twelve Black Friday price points on one video: submit twelve video-captions jobs with different cues at $0.20 each and a unique idempotency key per price.
- Subtitles in three languages: three caption jobs at $0.20 each
Each Sume standalone caption job costs $0.20 for video up to 60 seconds, and burns one set of cues. Three languages means three jobs, $0.60 in all.
Written by Sume