Avatar inline captions vs standalone video captions: what is billed
Inline captions on a Sume avatar video create no separate caption job. Standalone captions are a billed job on a public URL. When each one is right.
Inline captions on a Sume avatar talking video do not create a separate billed video-caption job. Standalone video captions are their own job on a public video URL you already have, at $0.20 per job for a clip of up to 60 seconds under the docs' current estimate. Use inline when you are rendering the avatar anyway. Use standalone when the video already exists.
Side by side
| Question | Inline on talking-video | Standalone video-captions |
|---|---|---|
| Input | Avatar script or video_inputs | Public HTTPS video_url |
| Settings | style, font, language, script_text | Same four, plus design, words, cues |
| Separate caption job | No | Yes |
| Duration limit | Rejects an estimate over 60 s | Priced for up to 60 s |
| If captions fail | Soft fail: video_url stays clean, captions.status=failed | Job fails with a typed error |
| Resource | Avatar video | GET /v1/video-captions/:id |
Soft fail is the main difference
A caption-stage failure inside the avatar job does not fail the job. You still get a clean primary video_url, and captions.status reads failed. Recaption that clean video with the standalone route, using style, language and, if needed, script_text.
Rules both share
- Korean text on
slam,punchortiktok-greenreturns400 caption_hangul_text_latin_style. - For Korean speech, select a Hangul style.
korean-adis the ad karaoke look. - Confirm the live price in
GET /v1/catalogbefore a batch.
A decision for a batch
Say you render ten avatar clips a week, each under 60 seconds. If you know the captions will be the same on every one, set them inline once in the request template. There is no second job to submit, poll or pay for.
If you later want to try another look on a finished clip, inline cannot help. Use standalone with the clip's public media.sume.com URL, or, if you made the first caption with the standalone route, send source_caption_id so Sume reuses the stored word timings and does not transcribe again.
If an avatar job returns captions.status of failed, the clean video_url is still valid. Do not resubmit the whole avatar job to fix captions. A resubmit pays for a new avatar render you do not need.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video aspect ratios: five choices, 720p only, 9:16 default
Sume Avatar 1.0 accepts 1:1, 3:4, 9:16, 4:3 and 16:9, defaults to 9:16, and renders 720p only. What each choice means for a vertical or landscape ad.
- Avatar package with captions and soundtrack: which video_url you get
In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.
- Avatar video with a product image: the premium is 1-3 cents a second
A product_image on a Sume Avatar 1.0 video adds $0.010 (standard), $0.013 (plus) or $0.030 (max) per second: 30 to 90 cents on a 30-second ad.
- Avatar video is ready but transcript_text is null: poll metadata again
A Sume avatar video can be resource_status ready while metadata.status is still processing. The file is usable now; the transcript and tags arrive later.
Written by Sume