script_text, words, cues or segments: which caption input to send
Sume captions take only one of script_text, words, cues, segments. Your pick decides whether speech-to-text runs and what happens on a silent clip.

Send exactly one of script_text, words, cues and segments to POST /v1/video-captions, or none. The Sume docs state you can send only one of them. The choice decides whether speech-to-text runs: with script_text it does, and Sume aligns your text to its word timings; with words, cues or segments it does not, and your text is burned at the times you give.
The four inputs
cues and segments are listed together in the docs as the phrase-level form.
| Input | Speech-to-text runs | Shape | Use it when |
|---|---|---|---|
| none | Yes | Only video_url | The transcript can be taken as spoken. |
| script_text | Yes, kept as the source of truth for time | One string | You have the exact script and the clip has speech. |
| words | No | Word-level, each with text, start, end in seconds | You have word times, for example from a prior transcript, or you need to correct text. |
| cues or segments | No | Phrase-level overlay cards with text, start, end | Silent clips and authored overlay text. |
Silent clips
Speech-to-captions works only when the clip has audible speech. On a silent clip the job fails with caption_no_speech, with next_action: use_overlay_captions. This is not a generic policy rejection. The fix is to send cues or segments with text, start and end, so Sume burns authored text without speech-to-text.
{
"video_url": "https://example.com/silent.mp4",
"cues": [
{ "text": "New this week", "start": 0.0, "end": 1.8 },
{ "text": "Free shipping", "start": 1.8, "end": 3.5 }
]
}When script_text can fail
script_text keeps the speech-to-text word timings as the source of truth and aligns your text to them. The alignment can fail with typed public job errors, which the video captions page lists. If your script and the speech differ a lot, expect failure, not a silent fallback. A language hint helps speech-to-text, and language never selects the style or the font.
Restyle without redoing the transcript
To burn the same video with another style, send source_caption_id and not video_url. Sume reuses the stored word timings, and you pass words only to correct text.
Sources
Related posts
More in Developers
- SDK waitForJob reads per minute: the 2-second floor, and 10 jobs
The Sume SDK polls a job at least every 2 seconds, and a longer next_poll_after_seconds wins. That is up to 30 reads a minute per job. A webhook removes them.
- See every webhook attempt for a Sume video job: webhook.delivery
Did Sume reach your endpoint? Read webhook_delivery on the job and the webhook.delivery events: statuses pending to exhausted, last_error and attempt count.
- Seedance 2.5 API 'coming soon' on ModelArk: what to call today
ByteDance says Seedance 2.5 API access is coming via ModelArk. Sume lists seedance-2.5 now: id, limits, price, and how to keep a later switch cheap.
- Seedance 2.5 green-screen editing: swap a background on Sume
ByteDance and fal describe green-screen editing for Seedance 2.5. Sume does not expose that mode, but Omni edit swaps a background with one video_url request.
Written by Sume