Avatar video succeeded but captions failed: what to do on Sume
On Sume a caption-stage failure is soft: the avatar job can still succeed with a clean video_url and captions.status=failed. How to check it and add captions.
If your Sume avatar video job succeeded but captions are missing, check captions.status in the result. A failure in the caption stage is a soft failure: the job can still succeed with a clean primary video_url and captions.status set to failed. Take that clean MP4 and run a standalone video caption job on it.
What the docs say
From Generate avatar video, read 2026-10-08:
- Captions are optional and burn into the clean final MP4 after generation.
- They use the spoken script or video_inputs text.
- Sume never adds captions to preview stills.
- Inline captions do not create a separate billed video-caption job.
- For inline captions, Sume rejects an estimated duration of more than 60 seconds.
Check, then recover
Read the result with GET /v1/jobs/{id}/result. If captions.status is failed, you still own a usable uncaptioned clip, so do not rerun the whole avatar job; that would be a second paid generation. Use the standalone route instead:
| Step | Call |
|---|---|
| Confirm | GET /v1/jobs/{id}/result and read captions.status |
| Re-caption | POST /v1/video-captions with the video_url of the clean MP4 |
| Read result | GET /v1/video-captions/{id} |
Mistakes that cause failures
Some failures are avoidable. A Korean script with style slam, punch or tiktok-green is rejected up front with 400 caption_hangul_text_latin_style, because those faces render Hangul as tofu; pick a Hangul style instead. A silent clip has no speech to transcribe, and the standalone job returns caption_no_speech with next_action use_overlay_captions. In that case send authored cues with text, start and end instead.
The standalone job accepts style, font, language and script_text. Sending script_text avoids relying on speech-to-text, which helps with names and product terms. See Video captions.
Preventing caption failures
Send a script whose language matches the caption style. For Korean, choose a Hangul style such as korean-ad, black-outline or pill-karaoke. For English, slam is the default. Keep the estimated duration at or below 60 seconds when using inline captions.
If speech is unclear, use the standalone caption route with script_text so the words come from your text and not from transcription. For brand names, this avoids misspellings.
Keep both versions of a video when you re-caption: the clean MP4 and the captioned one. A clean master lets you try different styles later without paying for a new avatar render.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video with a product image: +$0.01 to +$0.03 a second
Adding a product image to a Sume Avatar 1.0 video raises the per-second rate by $0.01 on standard, $0.013 on plus and $0.03 on max. Totals at 15, 30 and 60 s.
- Sume TTS voice then H3 Max lip sync: 40 seconds in 3 slices
H3 Max lip-sync accepts 5 to 14.8 s of audio. Split a 40 second TTS script into three sentence-based slices at $4.00 on 768p, and what to check.
- Can a Sume avatar speak my own recording? Voice types explained
Avatar 1.0 talking video takes a script or per-scene voice blocks of type text or silence. What the docs list, what they do not, and how to plan a pause.
- Will an AI video model lip-sync a voice-over added afterwards?
No: Sume's docs say video models do not lip-sync to generated speech or a later voice-over. A talking face needs a still plus audio, or an avatar video.
Written by Sume