Caption job 400: words, cues, segments, script_text are exclusive

The video captions API accepts only one wording source: words, cues or segments (which skip STT), or script_text (aligned onto STT). Sending two returns a 400.

4 min readSume
All posts

A caption job takes one source of wording. words, cues and segments all skip speech-to-text and name the same timed wording, so they are mutually exclusive with each other, and script_text is mutually exclusive with all three because there would be nothing left for it to align onto. Send two and the API returns a 400 with a message naming the conflict; the fix is to keep one.

Why the rule exists

The reasoning is in the request schema's own comments: silently picking one source would make the others' absence from the burn look like a bug, so the API refuses. hyperframes_composition is also exclusive with all four, because it is a Studio composition render rather than a karaoke overlay.

The video captions docs describe the inputs; the schema limits are 1 to 1200 words, 1 to 200 cues or segments, and script_text up to 8000 characters.

Which wording source to send (Sume docs and request schema, repo read 2026-10-05)
You haveSendSpeech-to-text runs?
Only the videovideo_urlYes
Video plus your scriptvideo_url and script_textYes, then aligned
Per-word timingswordsNo
Phrase timingscues or segmentsNo
An earlier caption jobsource_caption_id and a styleNo

Fixing a rejected request

Take the request that failed and drop everything but the source you trust. If you built words from an STT result and also left a script_text from an older attempt, remove the script. If you used both cues and segments because two tools wrote them, keep one.

The related error script_alignment_mismatch is different: it means script_text was accepted but could not be aligned to the speech, and the documented next action is to simplify the script or omit it.

Limits

punch and tiktok-green do not support design, and a Hangul style needs Hangul wording, which is checked against whichever source you send. A silent clip without cues returns caption_no_speech, whose next action is overlay captions. Each accepted standalone job is $0.20 under the current fixed estimate for videos up to 60 seconds; the live price comes from GET /v1/catalog.

Quick rules

Pick by what you already have, not by what looks the most complete. The mistake that triggers the 400 is usually a retry that adds a field to a previous attempt instead of replacing it, so build each request body from scratch.

  • Have a video and want the speech burned: send video_url alone.
  • Have a video and a clean script: add script_text.
  • Have timed words from another tool: send words and skip STT.
  • Have lines with timings and no per-word data: send cues; segments is an alias.
  • Want a new look on a finished job: send source_caption_id with a style, not a second STT pass.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume