Caption a silent Seedance 2.5 clip with authored cues

A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.

4 min readSume
All posts

If a Seedance 2.5 clip has no speech, Sume's caption job fails as caption_no_speech when it tries speech-to-text. Pass authored cues instead, each with text, start and end in seconds, and Sume burns exactly that copy at those times. A standalone caption job is a fixed $0.20 for videos up to 60 seconds, so captioning a 30-second clip costs far less than regenerating it.

Why generated text is the wrong tool

Generated video can garble small letters, and BytePlus's Dreamina Seedance 2.5 page, read 2026-10-04, says nothing about on-screen text accuracy; it lists 4-30 second output at 480p, 720p or 1080p. For anything a viewer must read exactly, such as a price, a name or a URL, burn the words on afterwards. The caption renderer draws the text you supply.

The request

The video captions docs say cues, segments, words and script_text are mutually exclusive, and that cues skips speech-to-text. The video_url must be a fetchable public HTTPS URL; localhost, private network, signed or private links, and provider task URLs are rejected.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-caption-001" \
  -d '{
    "video_url": "https://example.com/clips/trailer.mp4",
    "style": "punch",
    "cues": [
      { "text": "Learn it in a weekend", "start": 1.0, "end": 4.0 },
      { "text": "New course, open now", "start": 24.0, "end": 28.0 }
    ]
  }'

Timing the cues

Match cues to your beats. A 30-second clip with six five-second beats leaves room for a cue at the start of beat one and another near the end. Keep each cue short enough to read in its window, and leave gaps so the picture breathes.

Silent clip captioning on Sume (read 2026-10-04)
SituationApproachCost
Spoken words in the clipomit cues; speech-to-text$0.20 per job up to 60 s
Silent clipauthored cues$0.20 per job up to 60 s
Copy must match a scriptscript_text$0.20 per job up to 60 s
Regenerate to fix a typonew seedance-2.5 take, 30 s at 720p$17.33

Then deliver

Read the finished video_url from the job result as described in Jobs and results. If you need several captioned variants, restyle the first caption with source_caption_id instead of re-running speech-to-text, which the captions docs say keeps billing unchanged but reuses the timings.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume