Caption a silent Seedance 2.5 clip with authored cues
A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.

If a Seedance 2.5 clip has no speech, Sume's caption job fails as caption_no_speech when it tries speech-to-text. Pass authored cues instead, each with text, start and end in seconds, and Sume burns exactly that copy at those times. A standalone caption job is a fixed $0.20 for videos up to 60 seconds, so captioning a 30-second clip costs far less than regenerating it.
Why generated text is the wrong tool
Generated video can garble small letters, and BytePlus's Dreamina Seedance 2.5 page, read 2026-10-04, says nothing about on-screen text accuracy; it lists 4-30 second output at 480p, 720p or 1080p. For anything a viewer must read exactly, such as a price, a name or a URL, burn the words on afterwards. The caption renderer draws the text you supply.
The request
The video captions docs say cues, segments, words and script_text are mutually exclusive, and that cues skips speech-to-text. The video_url must be a fetchable public HTTPS URL; localhost, private network, signed or private links, and provider task URLs are rejected.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: silent-caption-001" \
-d '{
"video_url": "https://example.com/clips/trailer.mp4",
"style": "punch",
"cues": [
{ "text": "Learn it in a weekend", "start": 1.0, "end": 4.0 },
{ "text": "New course, open now", "start": 24.0, "end": 28.0 }
]
}'Timing the cues
Match cues to your beats. A 30-second clip with six five-second beats leaves room for a cue at the start of beat one and another near the end. Keep each cue short enough to read in its window, and leave gaps so the picture breathes.
| Situation | Approach | Cost |
|---|---|---|
| Spoken words in the clip | omit cues; speech-to-text | $0.20 per job up to 60 s |
| Silent clip | authored cues | $0.20 per job up to 60 s |
| Copy must match a script | script_text | $0.20 per job up to 60 s |
| Regenerate to fix a typo | new seedance-2.5 take, 30 s at 720p | $17.33 |
Then deliver
Read the finished video_url from the job result as described in Jobs and results. If you need several captioned variants, restyle the first caption with source_caption_id instead of re-running speech-to-text, which the captions docs say keeps billing unchanged but reuses the timings.
Sources
Related posts
More in Media tools
- Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.
- Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.
- Fix one wrong word in burned-in captions without a re-transcribe
One misheard word in a burned-in caption? Re-burn with source_caption_id and a words list instead of running speech-to-text again. Request body included.
- Cover or contain: fit a 16:9 clip into a vertical stack
Sume timeline compose takes video_fit cover, contain or stretch. Work out what each does to a 16:9 clip in the 1080x1920 default stack, then check one frame.
Written by Sume