Burn text onto a silent video: fixing caption_no_speech

A silent clip fails Sume's speech-to-captions path with caption_no_speech. Pass authored cues with text, start and end to burn overlay copy with no ASR.

5 min readSume
All posts

If your caption job fails with caption_no_speech, the clip has no audible speech and you did not give Sume any text. Pass cues (or segments) with text, start and end and Sume burns that copy without speech-to-text. The docs say the error carries next_action: use_overlay_captions, which is the same advice in machine form.

This is the right path for silent product clips, b-roll with on-screen titles, and music-only videos.

Why does a silent clip fail?

Speech-to-captions, which means script_text or speech-to-text when you omit both, needs audible speech to time the words. With none, there is nothing to align, so Sume returns a specific error and not a generic rejection. Authored cues do not need audio at all, because you supply the timing.

Caption input modes, read 2026-10-04
InputNeeds speechSkips STT
Nothing (default)YesNo
script_textYes, aligns your wording to STT timingsNo
wordsNoYes
cues or segmentsNoYes

What does the request look like?

The body has a public HTTPS video_url, an optional style and the cues. Times are in seconds, and the standalone job covers videos up to 60 seconds. script_text, words, cues and segments are mutually exclusive, so send only one.

import json, os, urllib.request

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("set SUME_API_KEY")
body = {
    "video_url": "https://example.com/silent-clip.mp4",
    "style": "slam",
    "cues": [
        {"text": "New this week", "start": 0.0, "end": 2.0},
        {"text": "Free shipping over $50", "start": 2.0, "end": 5.0},
    ],
}
req = urllib.request.Request(
    "https://api.sume.com/v1/video-captions",
    json.dumps(body).encode(),
    {"Authorization": f"Bearer {key}",
     "Content-Type": "application/json",
     "Idempotency-Key": "silent-cues-001"},
)
with urllib.request.urlopen(req, timeout=120) as r:
    print(json.load(r))

What else should you check?

A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate. Read Video captions for the full style list. For a worked cue request from a subtitle file, see the SRT post.

  • Latin text with a Latin style, and Hangul text with a Hangul style. Korean on a Latin style is rejected with caption_hangul_text_latin_style.
  • Whether the clip is really silent. A faint audio bed may be picked up as speech, so check with video inspect if the result surprises you.
  • Whether the cue text will read on a bright background; see the dim-a-clip post.

How do you time the cues?

With no speech to follow, you set the timing yourself. Watch the clip, note when each beat starts and ends, and leave about a half-second of hold after a short line so it can be read. Keep lines short, since a viewer has no audio cue to help them. Aim for cues that do not overlap, and make sure the last cue ends before the clip does.

A small check in code saves a failed job: confirm every end is greater than its start and every time is inside the first 60 seconds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume