AI video captions failed: is the video still usable?
On a Sume avatar video, a failed caption stage does not fail the job: the clean video_url still arrives with captions.status set to failed. What to do next.
Yes. When inline captions fail on a Sume avatar video, the caption stage soft-fails: the job can still succeed with a clean primary video_url and captions.status set to failed. You keep a usable video without burned-in text, and you decide whether to retry the captions on their own.
How do I tell that captions failed?
Poll the job like any other, then read the result. A completed job with a video does not mean captions were applied, so check captions.status in the result rather than only the job status.
The routes are GET /v1/jobs/{id}/status, /events and /result; completed results can include public media.sume.com video artifacts.
Which caption problems fail the job, and which soft-fail?
| Problem | What happens |
|---|---|
| Caption stage fails after generation | Job can succeed; clean video_url, captions.status=failed |
Korean script with slam, punch or tiktok-green | Rejected with 400 caption_hangul_text_latin_style |
| Estimated duration above 60 seconds | Rejected for inline captions |
| Silent clip on the standalone route | Fails as caption_no_speech |
What should I do when captions.status is failed?
- Use the clean
video_url; it is delivered whether or not the caption stage worked. - Retry captions on that public URL with the standalone Video captions route, which takes a
video_urlplusstyle,languageandscript_text. - Pass the exact spoken words as
script_textif the transcript was the problem. - Check the price of the standalone job first: inline captions do not create a separate billed caption job, but a standalone retry is a job of its own.
Does a soft-fail change what I keep?
The clean video_url is the primary output, so it is the file you publish or re-caption. The docs say the avatar job can still succeed, so treat captions.status=failed as a to-do rather than a lost render.
Log the caption status next to each job in your own system, so a batch of clips does not go out uncaptioned by accident.
Can I stop captions from failing in the first place?
- Send Hangul styles for Korean speech, so the request is not rejected before it starts.
- Keep the estimated duration at or under 60 seconds for inline captions.
- Give the exact spoken words as
script_textwhen the audio is noisy or the wording matters. - Remember that preview stills are never captioned, so a clean preview says nothing about the caption stage.
What if the clip has no speech?
A silent clip fails standalone speech-to-captions as caption_no_speech, with next_action: use_overlay_captions. For that case pass cues or segments with text, start and end to burn authored overlay copy without speech recognition.
On an avatar video, a silence scene has no speech by design, so captions come from the spoken scenes' text. See burn captions onto video with the API.
Sources
Related posts
More in Developers
- Bash for loop with curl: one API request per line
Loop over a file with while IFS= read -r, build each JSON body with jq --arg, send it with curl --fail-with-body, and pace it under the API's rate limit.
- Batch image generation API: how many images can run at once?
Sume queues extra generation jobs instead of rejecting them. The per-plan concurrency and queue table, the 429 queue_full case, and how to size an image batch.
- Batch video processing by API: one edit, many videos
Batch video processing by API is a loop: one edit job per file, each with its own idempotency key, collected by webhook. How it works on Sume, and costs.
- AI dubbing API: build the pipeline from STT, TTS and audio joins
Sume has no one-call dubbing route. Chain audio detach, STT 1.0, your translation, TTS 1.0 and Timeline audio into a dub. Routes, limits and the $0.01 steps.
Written by Sume