Silent gift-guide clip: caption_no_speech, send cues instead
A clip with no audible speech fails Sume video-captions with caption_no_speech. Send cues with text, start and end to burn gift-guide labels for $0.20.

A silent gift-guide clip fails Sume's video captions with caption_no_speech and next_action: use_overlay_captions, because speech-to-captions works only when the clip has audible speech. The fix is to send authored cues: a list of text, start and end in seconds that Sume burns exactly as written, with no speech-to-text. The job is $0.20 for a video up to 60 seconds.
This is the right path for a typical holiday gift guide: product shots, music or no sound, and a label for each item.
Pick the form that matches your clip
Captions have four ways to carry text. You can send only one per job.
| Field | What it is | Speech-to-text? |
|---|---|---|
| (none) | Speech captions from STT | Yes |
| script_text | Your script aligned to STT word times | Yes |
| words | Word-level authored text | No |
| cues or segments | Phrase-level overlay cards | No |
Label five gifts
Give each gift a card with a start and end. Leave a short gap between cards and make the last one end before the clip does. Keep the text short; a label of three to five words reads at a glance on a phone.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: gift-guide-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/silent.mp4",
"style": "black-outline",
"cues": [
{"text": "1. Wool scarf", "start": 0.0, "end": 2.5},
{"text": "2. Tea set", "start": 3.0, "end": 5.5}
]
}'Pitfalls
Do not send script_text for a silent clip. Sume keeps the speech-to-text word timings as the source of time for a script, so with no speech it cannot align, and an alignment failure shows as script_alignment_mismatch or script_alignment_failed; the recommended next action is to simplify the script or omit it.
For Korean cards on a silent clip, use a Hangul style such as black-outline; the Latin styles reject Hangul text. After the job, poll GET /v1/video-captions/:id and check the result against your list of gifts, because the text is burned exactly as you wrote it, including typos.
A 20 item guide split into four clips of five labels each is 4 x $0.20 = $0.80, and a single 60 second clip with 20 labels is $0.20. Choose by pacing, not cost.
- Silent clip plus speech captions = caption_no_speech.
- cues need text, start and end.
- Proofread; text is burned as written.
Timing a gift guide
A good rule is two and a half to three seconds per gift, with a half second gap. For a 30 second clip that gives about ten gifts. Write the cue list as data, not by hand in a JSON string: keep a sheet with gift name, start and end, and let a small script produce the cues array for each clip.
Because the job is billed per caption job, not per cue, packing all the labels into one request is the cheaper plan. A 60 second guide with 20 labels is one $0.20 job. Splitting it into 20 requests would be 20 x $0.20 = $4.00 and gives you 20 clips you do not want.
Make sure that your cue times match your edit. If you later trim or re-time the clip, the labels will drift, so caption last: trim and assemble first, then burn the labels on the final cut. A caption job should be the last step before upload, and the clean, uncaptioned master should be kept so you can reburn when the list changes.
Related posts
More in Media tools
- Skim a 20-minute video: video inspect seek fast vs precise
Use seek fast to sample 24 stills from a long video quickly, up to one keyframe interval early, and seek precise when a timestamp must match. Sume docs, read.
- Split one recording into many audio clips: detach once, then split
Twenty ranges from one talking-head video need one audio detach and one split job, not twenty cuts. Wav or mp3, overlap rules, and what each clip returns.
- Spotify and Amazon audio ads share one loudness target: -14 RMS
Spotify's video ad page and Amazon's audio ad page both ask for RMS at -14 dBFS and peak at -0.2 dBFS. Where Sume's gain_db helps and where you must measure.
- Square product photos to 9:16: timeline fit cover, contain or blur
Sume Timeline 1.0 renders 1080x1920 by default. Use video[].fit cover, contain, stretch or blur to place square product photos in a vertical holiday clip.
Written by Sume