Caption a silent clip with 3 timed cues: $0.20, no speech needed

A silent clip fails Sume video captions with caption_no_speech. Send cues with text, start and end to burn overlay text instead. One job, $0.20 up to 60 s.

4 min readSume
All posts

To caption a clip that has no speech, send cues instead of relying on speech-to-text. Each cue has text, start, and end in seconds, and Sume burns that text exactly at those times with no transcription. Without cues, a silent clip fails with caption_no_speech. The job costs $0.20 for a clip up to 60 seconds.

What the error looks like

Speech-to-captions only works when the clip has audible speech. On a silent clip the job fails with caption_no_speech, and the typed next_action is use_overlay_captions. This is not a policy rejection; it is the API telling you which request shape to switch to.

Three overlay cues

Cues are phrase-level overlay cards. You can send only one of script_text, words, cues, and segments per request (segments is an alias form of cues). Word-level words is also accepted if you need to place individual words.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/product.mp4",
    "style": "punch",
    "cues": [
      { "text": "New: matte black", "start": 0.5, "end": 2.5 },
      { "text": "Ships free", "start": 3, "end": 5 },
      { "text": "Today only", "start": 5.5, "end": 8 }
    ]
  }'

Cost and limits

Standalone caption jobs are a fixed estimate: $0.20 per accepted job, for videos of at most 60 seconds. Three cues or thirty make no difference to that number. The price is per job, not per second.

Cue form versus speech form (Sume docs, read 2026-10-09)
RequestUses speech-to-text?Works on silent clip?
no text fieldsYesNo: caption_no_speech
script_textYes, to time your textNo
wordsNoYes
cues or segmentsNoYes

Writing good cues

Keep each cue short enough to read in its window. Two seconds for three or four words is a comfortable pace for a phone screen. Cues should not overlap in time unless you want two cards at once, and the last end should be inside the clip; the docs do not state how out-of-range times are handled, so keep them within the video's duration.

If the same silent clip needs a different look, resend the request with a new style, or use source_caption_id to restyle without fetching the video again. Speech-to-text does not run in either case because your cues supply the text.

When the clip is not really silent

A clip with background music but no voice is a trap. Speech-to-text may find nothing and the job fails with caption_no_speech, even though the clip has an audio track. If you already know the clip has no speech, skip the failed attempt and go straight to cues. If you are unsure, the cost of finding out is a failed job; check the audio with video inspect first, whose probe.has_audio tells you whether there is an audio track at all, though not whether anyone speaks.

Constraints

video_url must be a public HTTPS video URL that Sume can fetch. Localhost, private networks, signed or private URLs, and provider task URLs are rejected, and SRT uploads are not supported: phrase-level text goes in cues. To change the look later without paying for another transcription, send source_caption_id instead of video_url; a restyle is still a render, so it is billed as one job. Poll GET /v1/jobs/:id/status or GET /v1/video-captions/:id for the burned video.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume