Avatar captions failed with script_alignment_mismatch: what to change

A pinned captions.script_text that does not align with the speech soft-fails the captions; the clean video stays. Match the spoken words, then re-caption.

4 min readSume
All posts

When captions.script_text does not line up with what the avatar actually said, Sume does not burn wrong words onto the video. The caption stage soft-fails with a typed public_reason such as script_alignment_mismatch, script_alignment_failed or empty_timed_words, and the job still succeeds with the clean MP4 as video_url. Fix it by making script_text match the spoken words, then caption the clean video with the standalone route.

Why a pinned script can fail

Per the Sume OpenAPI, inline captions derive wording from the spoken script by default. If you send script_text, you pin the on-screen copy: wording comes from your text and timing comes from speech-to-text, the same contract as standalone POST /v1/video-captions. When your text and the speech disagree too much to align, inline captions soft-fail with a typed reason and do not fall back to the transcribed words.

Without script_text, Sume derives caption copy from the spoken script and may fall back to speech-to-text wording on an alignment mismatch. So the strict behavior only applies to text you pinned yourself.

Caption outcomes on an avatar video (Sume OpenAPI, read 2026-10-05)
Situationcaptions.statusvideo_url
Captions on, no script_text, alignment finereadyCaptioned MP4
script_text pinned, alignsreadyCaptioned MP4
script_text pinned, does not alignfailed, with public_reasonClean MP4
Caption stage errorsfailed, caption_stage_failed possibleClean MP4
Captions omitted or enabled falsenullClean MP4

Typical causes and fixes

  • On-screen copy is a different sentence than the voice. Pin script_text only when it is the spoken script, word for word, or drop it.
  • Brand names, numbers, or symbols written differently than spoken, such as "2x" on screen but "twice" in the script. Write the spoken form in script_text.
  • Text that the avatar never speaks, such as a legal line. Put it in an authored overlay with cues on the standalone caption route instead.
  • A silent or near-silent clip. Speech-to-captions needs audible speech.

Recover without re-rendering the avatar

The avatar render succeeded, so do not pay for it again. Read the finished clean video_url and call the standalone caption route, which takes a public HTTPS video URL, style, language, and script_text. The docs note that inline captions do not create a separate billed caption job, so the standalone route is a separate job to plan for. Unlike inline captions, standalone captions still hard-fail on an alignment problem, so fix the text first.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recaption-avatar-video-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/your-clean-video.mp4",
    "style": "slam",
    "language": "en",
    "script_text": "The exact words the avatar speaks."
  }'

What not to do

Do not use script_text to change the wording of a caption for effect. It is a contract with the audio, not a copy editor. For a different on-screen line, use authored cues. Also remember the Korean rule: Latin-only styles on Hangul text are rejected with caption_hangul_text_latin_style, which is a separate error from alignment.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume