Caption job 400: words, cues, segments, script_text are exclusive
The video captions API accepts only one wording source: words, cues or segments (which skip STT), or script_text (aligned onto STT). Sending two returns a 400.

A caption job takes one source of wording. words, cues and segments all skip speech-to-text and name the same timed wording, so they are mutually exclusive with each other, and script_text is mutually exclusive with all three because there would be nothing left for it to align onto. Send two and the API returns a 400 with a message naming the conflict; the fix is to keep one.
Why the rule exists
The reasoning is in the request schema's own comments: silently picking one source would make the others' absence from the burn look like a bug, so the API refuses. hyperframes_composition is also exclusive with all four, because it is a Studio composition render rather than a karaoke overlay.
The video captions docs describe the inputs; the schema limits are 1 to 1200 words, 1 to 200 cues or segments, and script_text up to 8000 characters.
| You have | Send | Speech-to-text runs? |
|---|---|---|
| Only the video | video_url | Yes |
| Video plus your script | video_url and script_text | Yes, then aligned |
| Per-word timings | words | No |
| Phrase timings | cues or segments | No |
| An earlier caption job | source_caption_id and a style | No |
Fixing a rejected request
Take the request that failed and drop everything but the source you trust. If you built words from an STT result and also left a script_text from an older attempt, remove the script. If you used both cues and segments because two tools wrote them, keep one.
The related error script_alignment_mismatch is different: it means script_text was accepted but could not be aligned to the speech, and the documented next action is to simplify the script or omit it.
Limits
punch and tiktok-green do not support design, and a Hangul style needs Hangul wording, which is checked against whichever source you send. A silent clip without cues returns caption_no_speech, whose next action is overlay captions. Each accepted standalone job is $0.20 under the current fixed estimate for videos up to 60 seconds; the live price comes from GET /v1/catalog.
Quick rules
Pick by what you already have, not by what looks the most complete. The mistake that triggers the 400 is usually a retry that adds a field to a previous attempt instead of replacing it, so build each request body from scratch.
- Have a video and want the speech burned: send
video_urlalone. - Have a video and a clean script: add
script_text. - Have timed words from another tool: send
wordsand skip STT. - Have lines with timings and no per-word data: send
cues;segmentsis an alias. - Want a new look on a finished job: send
source_caption_idwith a style, not a second STT pass.
Sources
Related posts
More in Developers
- Chain a 30-second render, trim and captions: three jobs, three keys
Generate, trim and caption an AI clip through the Sume API as three separate jobs, each with its own Idempotency-Key, status poll and result read.
- Check a callback_url is public HTTPS before you submit a Sume video
Sume's callback_url must be public HTTPS. A Python pre-check rejects http, embedded credentials, unresolvable hosts and private addresses before you pay.
- Check a Short before you upload: duration, size and audio in one probe
Before posting a 9:16 file, probe it. Sume video inspect returns the probe and up to 24 stills for free. Compare duration with YouTube's 3-minute Shorts limit.
- Check a video request against the Sume catalog before the 400
Sume returns 400 unsupported_capability for unlisted values. A Python pre-flight reads GET /v1/videos/models and flags a bad duration or ratio first.
Written by Sume