caption_no_speech: burn authored cues onto a silent clip, no ASR

A silent clip makes speech-to-captions fail with caption_no_speech. Send cues (text, start, end) and Sume burns your text at those times without speech-to-text.

4 min readSume
All posts

If a clip has no audible speech, the caption job fails with caption_no_speech and next_action: use_overlay_captions. The fix is to send cues (or segments), each with text, start and end in seconds. Sume then burns your text at those times and does not run speech-to-text.

Which input does what

You can send only one of the four forms in a request.

Caption inputs, per the video-captions docs read 2026-10-09
FieldGranularityRuns speech-to-text?Use for
(none)words from STTyesa clip with speech
script_textaligned to STT timingsyesa known script
wordsword-level, yoursnocorrected transcript
cues / segmentsphrase cards, yoursnosilent clips, overlay text

A silent product loop

A generated product loop with no voice is the usual case. Write three cards: 0 to 2.5 s, 2.5 to 5 s, 5 to 8 s. Keep each card short; the phrasing controls (max_words, max_chars) apply to speech-derived words, so for cues you set the line length yourself.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cues-001" \
  -d '{"video_url":"https://media.sume.com/artifacts/example/loop.mp4","style":"punch","cues":[{"text":"New drop","start":0,"end":2.5},{"text":"Ships Friday","start":2.5,"end":5},{"text":"Link in bio","start":5,"end":8}]}'

Price and checks

The job is the standalone caption job: $0.20 for a clip of up to 60 seconds under the current estimate (live price in the catalog). Because there is no transcription step, you pay for the render only in effect, but the docs state a single fixed price, so budget $0.20.

Before you submit

  • Make sure end is after start and inside the clip length.
  • Korean text on slam, punch or tiktok-green returns caption_hangul_text_latin_style; pick a Hangul style.
  • The video_url must be public HTTPS.

Timing guidance

Keep cues at least 1.5 seconds long so a viewer can read them, and avoid overlap, since cards are placed one at a time. For an 8-second loop, three cards of 2.5, 2.5 and 3 seconds work. If you later add a voiceover, re-run with speech-to-captions on the new file; the cue form is only for text you wrote. A failed job returns a typed error with the next action, which means you can branch your pipeline on caption_no_speech rather than on message text. Every media job follows the same lifecycle: submit with an Idempotency-Key, receive a job, poll GET /v1/jobs/:id/status until it is ready, then read GET /v1/jobs/:id/result. A retry with the same key does not queue a second job, so a network error during submit never doubles a charge.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume