caption_no_speech: burn authored cues onto a silent clip, no ASR
A silent clip makes speech-to-captions fail with caption_no_speech. Send cues (text, start, end) and Sume burns your text at those times without speech-to-text.

If a clip has no audible speech, the caption job fails with caption_no_speech and next_action: use_overlay_captions. The fix is to send cues (or segments), each with text, start and end in seconds. Sume then burns your text at those times and does not run speech-to-text.
Which input does what
You can send only one of the four forms in a request.
| Field | Granularity | Runs speech-to-text? | Use for |
|---|---|---|---|
| (none) | words from STT | yes | a clip with speech |
script_text | aligned to STT timings | yes | a known script |
words | word-level, yours | no | corrected transcript |
cues / segments | phrase cards, yours | no | silent clips, overlay text |
A silent product loop
A generated product loop with no voice is the usual case. Write three cards: 0 to 2.5 s, 2.5 to 5 s, 5 to 8 s. Keep each card short; the phrasing controls (max_words, max_chars) apply to speech-derived words, so for cues you set the line length yourself.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: cues-001" \
-d '{"video_url":"https://media.sume.com/artifacts/example/loop.mp4","style":"punch","cues":[{"text":"New drop","start":0,"end":2.5},{"text":"Ships Friday","start":2.5,"end":5},{"text":"Link in bio","start":5,"end":8}]}'Price and checks
The job is the standalone caption job: $0.20 for a clip of up to 60 seconds under the current estimate (live price in the catalog). Because there is no transcription step, you pay for the render only in effect, but the docs state a single fixed price, so budget $0.20.
Before you submit
- Make sure
endis afterstartand inside the clip length. - Korean text on
slam,punchortiktok-greenreturnscaption_hangul_text_latin_style; pick a Hangul style. - The
video_urlmust be public HTTPS.
Timing guidance
Keep cues at least 1.5 seconds long so a viewer can read them, and avoid overlap, since cards are placed one at a time. For an 8-second loop, three cards of 2.5, 2.5 and 3 seconds work. If you later add a voiceover, re-run with speech-to-captions on the new file; the cue form is only for text you wrote. A failed job returns a typed error with the next action, which means you can branch your pipeline on caption_no_speech rather than on message text. Every media job follows the same lifecycle: submit with an Idempotency-Key, receive a job, poll GET /v1/jobs/:id/status until it is ready, then read GET /v1/jobs/:id/result. A retry with the same key does not queue a second job, so a network error during submit never doubles a charge.
Sources
Related posts
More in Media tools
- Caption a silent clip with 3 timed cues: $0.20, no speech needed
A silent clip fails Sume video captions with caption_no_speech. Send cues with text, start and end to burn overlay text instead. One job, $0.20 up to 60 s.
- Check pause-cut seams with video frames: 24 stills cover 12 seams
Video frames returns up to 24 stills per call. Two stills per seam covers 12 seams per call, so a 14-cut Reel needs 2 calls to see every join.
- Concat 20 voice lines into one wav with timeline audio for 1 cent
Timeline audio concat joins up to 20 Sume-hosted audio parts into one gapless wav for a flat $0.01, and returns segment offsets to re-base your video slots.
- Conform a trim to 540x960 at 30 fps for TikTok non-Spark ads
TikTok's non-Spark ad spec sets 540x960 as the 9:16 minimum. One video-trim job with output 540x960 at 30 fps costs $0.02. Bitrate is not a field.
Written by Sume