script_alignment_mismatch: fixing a caption job with script_text
When script_text does not match the speech, a Sume caption job fails with script_alignment_mismatch or script_alignment_failed. Simplify the script or omit it.

script_alignment_mismatch means the script_text you sent does not line up with what speech-to-text heard. Sume keeps the speech-to-text word timings as the source of truth for time and tries to align your script to them. If it cannot, the job fails with script_alignment_mismatch or script_alignment_failed, and the recommended next action is simplify_script_text_or_omit.
Why alignment fails
Sume does not retime your script by guessing. It places your words onto the heard words, so large differences break the match. The docs do not publish a numeric tolerance, so do not plan around one; treat any substantial rewrite as a risk.
- The speaker ad-libbed or skipped lines that are in your script.
- The script has stage directions, speaker labels or numerals written differently from how they were spoken.
- The recording is a different take than the script.
- The clip has more than one language, or a language other than the one you expected.
Your three options
| Option | What you send | Effect |
|---|---|---|
| Simplify | Shorter script_text with plain words that were spoken | Alignment can succeed |
| Omit | No script_text | Sume burns the speech-to-text words |
| Take over | words with start and end, or cues | No speech-to-text; Sume burns your text at those times |
Choosing between them
Use the simplified script when the goal is correct brand spelling, such as a product name that speech-to-text writes wrong. Keep the script to the lines that were actually said, in the order they were said.
Omit the script when accuracy matters less than speed. Use words when you have timings from another tool, for example from video inspect with transcribe: true and sentence segmentation, and you only need Sume to burn them in.
Request that includes a script
The docs' own example sends script_text with the punch style.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-caption-script-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "punch",
"script_text": "Say hello to the Sume developer platform."
}'Diagnosing a failed alignment
Read the job's error code, not only the message. Both alignment codes come back as typed public job errors with the same recommended action, so one handler can cover them. A good handler tries once with a trimmed script, then once with no script, and finally logs the clip for a human if brand spelling is critical.
If you only need a handful of names spelled right, a short script of those exact phrases is more likely to align than a full paragraph, but the docs only promise that simplifying is the recommended next step, not that it always succeeds.
- Do not send
script_texttogether withwords,cuesorsegments; only one of the four is allowed per request. - Burning the speech-to-text words needs no script at all.
Sources
Related posts
More in Media tools
- Six 5-second reference clips: $2.28 on H3 to $17.34 on Seedance 2.5
Cost of six 5-second 720p reference-to-video clips on Sume by model: MiniMax H3 $2.28 at 768p, Wan 3.0 $3.78, Seedance 2.0 Fast $9.12, Seedance 2.5 $17.34.
- Timeline soundtrack duck_db: 0 to 20 dB, needs a real audio spine
Sume Timeline soundtrack.duck_db accepts 0 to 20 and returns duck_requires_audio_spine on silence. Use it with a voiceover spine; gain_db covers a static level.
- STT-ready audio: detach as 16000 Hz mono wav in one call
Sume audio detach can output the STT shape directly: sample_rate 16000 with channels mono, in sample-exact wav, for $0.01 per job.
- Three 61-second renders cost $0.60; one 183-second render costs $0.40
Timeline 1.0 rounds each render up to a whole minute at $0.10. Three 61 s renders bill 6 minutes; one 183 s render bills 4. A YouTube Shorts length check too.
Written by Sume