Avatar captions failed with script_alignment_mismatch: what to change
A pinned captions.script_text that does not align with the speech soft-fails the captions; the clean video stays. Match the spoken words, then re-caption.
When captions.script_text does not line up with what the avatar actually said, Sume does not burn wrong words onto the video. The caption stage soft-fails with a typed public_reason such as script_alignment_mismatch, script_alignment_failed or empty_timed_words, and the job still succeeds with the clean MP4 as video_url. Fix it by making script_text match the spoken words, then caption the clean video with the standalone route.
Why a pinned script can fail
Per the Sume OpenAPI, inline captions derive wording from the spoken script by default. If you send script_text, you pin the on-screen copy: wording comes from your text and timing comes from speech-to-text, the same contract as standalone POST /v1/video-captions. When your text and the speech disagree too much to align, inline captions soft-fail with a typed reason and do not fall back to the transcribed words.
Without script_text, Sume derives caption copy from the spoken script and may fall back to speech-to-text wording on an alignment mismatch. So the strict behavior only applies to text you pinned yourself.
| Situation | captions.status | video_url |
|---|---|---|
| Captions on, no script_text, alignment fine | ready | Captioned MP4 |
| script_text pinned, aligns | ready | Captioned MP4 |
| script_text pinned, does not align | failed, with public_reason | Clean MP4 |
| Caption stage errors | failed, caption_stage_failed possible | Clean MP4 |
| Captions omitted or enabled false | null | Clean MP4 |
Typical causes and fixes
- On-screen copy is a different sentence than the voice. Pin
script_textonly when it is the spoken script, word for word, or drop it. - Brand names, numbers, or symbols written differently than spoken, such as "2x" on screen but "twice" in the script. Write the spoken form in
script_text. - Text that the avatar never speaks, such as a legal line. Put it in an authored overlay with
cueson the standalone caption route instead. - A silent or near-silent clip. Speech-to-captions needs audible speech.
Recover without re-rendering the avatar
The avatar render succeeded, so do not pay for it again. Read the finished clean video_url and call the standalone caption route, which takes a public HTTPS video URL, style, language, and script_text. The docs note that inline captions do not create a separate billed caption job, so the standalone route is a separate job to plan for. Unlike inline captions, standalone captions still hard-fail on an alignment problem, so fix the text first.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recaption-avatar-video-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/your-clean-video.mp4",
"style": "slam",
"language": "en",
"script_text": "The exact words the avatar speaks."
}'What not to do
Do not use script_text to change the wording of a caption for effect. It is a contract with the audio, not a copy editor. For a different on-screen line, use authored cues. Also remember the Korean rule: Latin-only styles on Hangul text are rejected with caption_hangul_text_latin_style, which is a separate error from alignment.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar Face Swap (Beta) for ad variants: 4 to 15 second source clips
Sume's Avatar Face Swap Beta puts a ready avatar face on a public source video. What it requires, what it does not take, and how to poll the job.
- Avatar inline captions vs standalone video captions: what is billed
Inline captions on a Sume avatar video create no separate caption job. Standalone captions are a billed job on a public URL. When each one is right.
- Avatar video aspect ratios: five choices, 720p only, 9:16 default
Sume Avatar 1.0 accepts 1:1, 3:4, 9:16, 4:3 and 16:9, defaults to 9:16, and renders 720p only. What each choice means for a vertical or landscape ad.
- Avatar package with captions and soundtrack: which video_url you get
In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.
Written by Sume