Burn text onto a silent video: fixing caption_no_speech
A silent clip fails Sume's speech-to-captions path with caption_no_speech. Pass authored cues with text, start and end to burn overlay copy with no ASR.

If your caption job fails with caption_no_speech, the clip has no audible speech and you did not give Sume any text. Pass cues (or segments) with text, start and end and Sume burns that copy without speech-to-text. The docs say the error carries next_action: use_overlay_captions, which is the same advice in machine form.
This is the right path for silent product clips, b-roll with on-screen titles, and music-only videos.
Why does a silent clip fail?
Speech-to-captions, which means script_text or speech-to-text when you omit both, needs audible speech to time the words. With none, there is nothing to align, so Sume returns a specific error and not a generic rejection. Authored cues do not need audio at all, because you supply the timing.
| Input | Needs speech | Skips STT |
|---|---|---|
| Nothing (default) | Yes | No |
script_text | Yes, aligns your wording to STT timings | No |
words | No | Yes |
cues or segments | No | Yes |
What does the request look like?
The body has a public HTTPS video_url, an optional style and the cues. Times are in seconds, and the standalone job covers videos up to 60 seconds. script_text, words, cues and segments are mutually exclusive, so send only one.
import json, os, urllib.request
key = os.environ.get("SUME_API_KEY", "")
if not key:
raise SystemExit("set SUME_API_KEY")
body = {
"video_url": "https://example.com/silent-clip.mp4",
"style": "slam",
"cues": [
{"text": "New this week", "start": 0.0, "end": 2.0},
{"text": "Free shipping over $50", "start": 2.0, "end": 5.0},
],
}
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
json.dumps(body).encode(),
{"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
"Idempotency-Key": "silent-cues-001"},
)
with urllib.request.urlopen(req, timeout=120) as r:
print(json.load(r))What else should you check?
A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate. Read Video captions for the full style list. For a worked cue request from a subtitle file, see the SRT post.
- Latin text with a Latin style, and Hangul text with a Hangul style. Korean on a Latin style is rejected with
caption_hangul_text_latin_style. - Whether the clip is really silent. A faint audio bed may be picked up as speech, so check with video inspect if the result surprises you.
- Whether the cue text will read on a bright background; see the dim-a-clip post.
How do you time the cues?
With no speech to follow, you set the timing yourself. Watch the clip, note when each beat starts and ends, and leave about a half-second of hold after a short line so it can be read. Keep lines short, since a viewer has no audio cue to help them. Aim for cues that do not overlap, and make sure the last cue ends before the clip does.
A small check in code saves a failed job: confirm every end is greater than its start and every time is inside the first 60 seconds.
Sources
Related posts
More in Developers
- C2PA 2.2: file types that can carry credentials vs Sume outputs
C2PA 2.2 manifests can be embedded in JPEG, PNG, WebP, SVG, MP4, MOV and more. How that list lines up with the formats Sume image and video jobs return.
- C2PA 2.2 in plain terms: manifests, hard and soft bindings
C2PA 2.2 describes signed manifests with hash-based hard bindings and fingerprint or watermark soft bindings. What that implies after a re-encode or trim.
- Canary 5-10% of image jobs to a new Sume image model in Python
Route a stable slice of image jobs to a new model by hashing the job key, log the model and cost, and widen the slice only if review scores hold.
- Cancel a Sume Omni job: only before it starts, then 409
POST /v1/jobs/{id}/cancel works while a Sume job is queued. Once generation starts you get 409 job_generation_already_started and the clip is still billed.
Written by Sume