caption_no_speech: caption a silent clip with cues ($0.20)
Sume's caption job fails with caption_no_speech on silent clips. Send cues with text, start and end in seconds to burn authored text without speech-to-text.

The error and the fix
A caption job that relies on speech fails with caption_no_speech if the clip has no audible speech. The error carries next_action use_overlay_captions and is not a generic policy rejection. The fix is to send cues, or segments, each with text, start and end in seconds; Sume burns that text at those times and does not run speech-to-text.
Input forms
Only one of script_text, words, cues and segments may be sent in a request.
| Field | Granularity | Speech-to-text runs? |
|---|---|---|
| (none) | Heard words | Yes |
| script_text | Your script aligned to heard timing | Yes |
| words | Word-level | No |
| cues or segments | Phrase-level overlay cards | No |
Example body
A 20-second silent product loop with three text cards would send three cues, for example start 0 end 6 for the first card. Each unit is seconds, not frames.
The job price is the same $0.20 for a clip up to 60 seconds, so silent overlays are not cheaper than spoken captions.
Planning around the error
When a pipeline mixes spoken and silent clips, call video inspect first with frames false and read probe.has_audio, which is free, and route silent clips to cues. A video with music but no speech also lacks audible speech for transcription purposes, so treat has_audio as necessary but not sufficient.
Reference: Video captions and Video inspect.
Timing the cues
Cues use seconds from the start of the clip. Aim for a minimum on-screen time that a viewer can read, roughly the length of the text read aloud once, and avoid overlapping cues. Overlap makes both cards compete for the same space.
Check probe.has_audio and listen to the track before assuming a clip is silent.
Handling it in a pipeline
Treat caption_no_speech as a routing signal, not a failure. Catch the code, read next_action, and resubmit with cues generated from your own copy deck. Because the price is the same, no budget logic needs to change. Log which clips took the overlay path so editors know where text was authored by hand.
Sources
More in Developers
- Cartesia Line SDK hosting ends Dec 1 2026: does batch TTS change?
Cartesia's changelog says Line SDK agent hosting ends Dec 1, 2026 and agent LLM charges began Oct 1. What that means, and what it does not for batch TTS jobs.
- Cartesia sonic-3-latest and sonic-3-preview aliases: fix model strings
Cartesia lists sonic-3-latest and sonic-3-preview as deprecated aliases. What each maps to, which ids Sume's TTS Router accepts, and a safe migration.
- Celery task for a Sume video: task id as the Idempotency-Key
One Celery task submits POST /v1/videos with its own task id as the Idempotency-Key and re-queues itself every 30 seconds until the job ends. No double bills.
- Cheapest Sume image model for a 21:9 banner from a reference photo
Qwen Image at $0.025 is the cheapest row listing 21:9 with reference input. Flux 2 Pro is next at $0.0375. A short script finds it live.
Written by Sume