Muted product loop in four languages: burn captions from cues
A silent clip fails speech-to-captions as caption_no_speech. Pass cues with text, start and end to burn four translations at $0.20 a job.

A silent product loop has no speech for transcription to hear, so the standard caption job fails with caption_no_speech and the next action use_overlay_captions. The fix on Sume is to pass cues (or segments), each with text, start and end in seconds; that skips speech-to-text and burns exactly your copy at exactly those times. For four languages you submit four jobs at $0.20 each for clips up to 60 seconds, so $0.80 in total.
Cues versus the other inputs
The video captions page defines four ways to supply wording. Only one of them is allowed per request: script_text, words, cues and segments are mutually exclusive.
| Input | Needs speech? | Use when |
|---|---|---|
| Nothing (speech-to-text) | Yes | Talking-head clip, no script |
script_text | Yes | Speech exists and you want exact spelling |
words | No | You already have word-level timings |
cues or segments | No | Silent clips and authored overlay copy |
A four-language set
Keep the timing identical across languages and change only the text. If the loop is 8 seconds with two messages, the same start and end apply to each translation.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: loop-es-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/loop.mp4",
"style": "slam",
"cues": [
{ "text": "Nuevo sabor", "start": 0.5, "end": 3.5 },
{ "text": "Ya a la venta", "start": 4.0, "end": 7.5 }
]
}'Checks before you ship
Translations change length. A cue that fits in English can wrap badly in German, so look at each render. The design field can edit the placement, type size and phrasing limits for a request, as described in the video captions docs, and numbers outside the documented range are refused at request time rather than rendering wrong. design is not supported on punch or tiktok-green.
Script matters too. Hangul text sent to the Latin styles slam, punch and tiktok-green is rejected with caption_hangul_text_latin_style, so a Korean version needs one of the Hangul-safe styles such as black-outline. Write the right style per language rather than one style for all four.
Batching
Use one Idempotency-Key per language and clip so a retry cannot double-charge. If you submit with a webhook, the webhooks page says only terminal events are sent, so keep status polling as a backup. When the loop changes, a new cue file means a new burn; there is no way to edit burned pixels.
Sources
Related posts
More in Use cases
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Near-duplicate AI images: perceptual-hash dedupe before human review
Four-up image batches often contain near twins. Use a 64-bit difference hash in Pillow to drop duplicates before a person reviews them. Python for Sume results.
- Neighborhood holiday lights clip for agents from three street photos
A real estate agent's neighborhood lights tour from three of your own street photos: three 4-second Wan 3.0 shots at 480p and a join, about $0.85 on Sume.
- Netflix Ads: 10-75 s at 16:9 1920x1080, render 16:9 native
Netflix Ads video specs: 16:9 at 1920x1080, 10 to 75 seconds, H.264. Render 16:9 natively on Sume and sequence clips for longer spots.
Written by Sume