Caption a silent AI video: fixing caption_no_speech

A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.

4 min readSume
All posts

A video with no audible speech fails POST /v1/video-captions with caption_no_speech, because the default path runs speech-to-text. Pass cues (or segments) with text, start and end in seconds instead. Sume then burns exactly that copy at those times and skips speech-to-text.

Why a silent clip fails

Standalone captions take a public HTTPS video_url. With script_text, or with nothing, Sume needs audible speech to time the words. A silent clip returns a typed failure caption_no_speech with next_action: use_overlay_captions. It is not a generic policy rejection, and the fix is in the request, not the video.

Send cues instead

cues and segments are phrase-level overlay cards. words is word-level. script_text, words, cues and segments are mutually exclusive, so send one.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-caption-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/silent.mp4",
    "style": "slam",
    "cues": [
      {"text": "New drop", "start": 0.2, "end": 1.8},
      {"text": "Out this Friday", "start": 2.0, "end": 4.0}
    ]
  }'

Style and language choices

Omit style and the wording decides: slam for Latin text, black-outline for Korean. You can also name slam, punch, tiktok-green or korean-ad. Korean copy on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style, because those faces have no Hangul glyphs; pick a Hangul style for Korean cues.

design overrides colors, typography, placement, phrasing and motion for one request, but it is not supported on punch or tiktok-green.

Cost, limits and polling

Each accepted standalone caption job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog. The video_url must be a fetchable public HTTPS URL. Localhost, private-network, signed or private URLs are rejected, and SRT uploads are not supported; send phrase text as cues instead.

Poll GET /v1/jobs/{id}/status and read GET /v1/video-captions/{id} when ready. To restyle the same video later, pass source_caption_id instead of video_url; billing is unchanged because a restyle is still a render.

Caption input choices (read 2026-10-03)
InputNeeds speechUse when
none or script_textYesThe clip has a voice
wordsNoYou have word timings
cues or segmentsNoThe clip is silent or you author the copy

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume