Avatar UGC ad: captions.status failed but the video is fine, now what
Inline captions on a Sume avatar video soft-fail: the job can succeed with a clean video_url and captions.status failed. Re-caption it standalone for $0.20.
If an avatar video job completes but captions.status reads failed, the ad itself is fine. Sume's docs say caption-stage failures soft-fail: the job can still succeed with a clean primary video_url. You do not need to re-run the avatar video. Take that clean file and send it to the standalone video captions endpoint.
This matters in a seasonal batch because a retry of the whole job would pay for the avatar render a second time. Re-captioning only pays the caption step.
How do inline captions work on an avatar video?
On POST /v1/avatar-1.0/talking-video you can add a captions object. After generation, Sume burns the style into the clean final MP4 using the spoken script or the video_inputs text. Preview stills are never captioned.
The object takes the same four knobs as the standalone endpoint: style, optional font, a language hint and script_text. The default style is slam. Estimated durations above 60 seconds are rejected for inline captions, and Avatar Video itself accepts 4 to 60 seconds.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: holiday-ugc-sku-2041" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"script": "I wrapped three of these for gifts and they looked great.",
"captions": { "enabled": true, "style": "slam", "language": "auto" }
}'What does a soft-failed caption stage look like?
You read the finished job at GET /v1/jobs/:id/result as usual. A soft failure leaves the job completed with a usable video_url, and the captions block reports status: failed. Check that field in your pipeline, not only the job status, or a captionless clip will slip through to the ad account.
The docs do not list a specific error code for this stage, so we will not invent one. If you need more detail, the job events endpoint (GET /v1/jobs/:id/events, see Jobs and results) is the place to look, and keep the job id if you contact support.
How do you recover without paying for the avatar again?
Send the clean video_url from the result to POST /v1/video-captions with the style and language you wanted. The standalone job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate; confirm the live number in GET /v1/catalog.
Inline avatar captions are a separate add-on on the avatar-video estimate and do not create a video_caption resource. That means you have no source_caption_id to reuse after an inline failure, so the standalone call will run speech-to-text from the video. If you have the approved script, pass it as script_text to keep the burned wording on-script.
| Path | Creates a video_caption resource | Failure behaviour | Billing |
|---|---|---|---|
| Inline captions on talking-video | No | Soft-fails; clean video_url stays, captions.status is failed | Add-on on the avatar-video estimate |
| Standalone POST /v1/video-captions | Yes | Typed job errors such as caption_no_speech | $0.20 up to 60 s, fixed estimate |
What should a holiday batch do about it?
Treat captions as a checked output of every job. In your poller, branch on three states: completed with captions, completed with captions.status of failed, and failed. Only the second one gets the cheap recovery above.
One more guard: a Korean script with slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style at request time, before any billing. Pick a Hangul style for Korean speech. For wording that must match legal copy exactly, our suggestion is to pass script_text rather than rely on transcription alone, though that adds the alignment errors script_alignment_mismatch and script_alignment_failed as a possible failure.
Sources
Related posts
More in Use cases
- Black Friday audio ad read: a TTS voice over a music bed
Produce a 30-second Black Friday ad read: write to character count, generate with Sume TTS, add a short music bed and mix. Costs and limits included.
- Build a 3-minute YouTube Short from clips with Sume Timeline
Stitch several clips into one Short of up to 180 seconds with Timeline 1.0: fades, a looped music bed with ducking, and the $0.30 render for three minutes.
- Bulk run says completed: re-queue only the failed holiday SKUs
A Sume bulk queue is completed once every item is terminal, not once all succeed. Read counts.failed, then re-queue only those under a new idempotency key.
- Character turnaround sheet by API: front, profile, back views
Make a turnaround sheet by API: render the front view, then send it as a reference for profile and back. Google's 360-view method, and the Sume request.
Written by Sume