Burn feature text onto silent product clips with caption cues
A silent product clip fails speech captions. Pass your own cues (text, start, end) to Sume's video-captions endpoint and burn feature text on for TikTok Shop.

To put feature text on a silent product clip, call POST /v1/video-captions with the clip's public HTTPS URL and a cues array of text, start and end in seconds. Sume skips speech-to-text and burns exactly your copy at those times. Without cues, a clip with no speech fails with caption_no_speech.
Short, readable on-screen text matters when video is the storefront. Best Buy's TikTok Shop debut in late October 2026 leans on shoppable creator videos (read 2026-10-04), so the clip itself has to say what the product is without sound.
Why do silent clips fail the default path?
Speech-to-captions, whether from script_text or from speech-to-text when it is omitted, needs audible speech. The docs mark a silent clip as caption_no_speech with next_action use_overlay_captions. That is a pointer to the cues path, not a dead end.
What are the request options?
Only video_url is required.
| Field | Meaning |
|---|---|
| video_url | Public HTTPS video URL (required) |
| cues or segments | Authored overlay cards with text, start, end; skips speech-to-text |
| script_text, words, cues, segments | Mutually exclusive |
| style | slam, punch, tiktok-green for Latin text; Hangul styles for Korean |
| Price | $0.20 per accepted job |
What does the call look like?
Keep each card short, and keep end times inside the clip. The call returns a job; read it from the job endpoints.
import os, requests
r = requests.post(
"https://api.sume.com/v1/video-captions",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json={
"video_url": "https://example.com/earbuds-silent.mp4",
"style": "slam",
"cues": [
{"text": "Fits in one hand", "start": 0.5, "end": 2.5},
{"text": "Charging case included", "start": 3.0, "end": 5.5},
],
},
timeout=60,
)
print(r.status_code, r.text[:300])
What text should you burn?
Only facts you can show or stand behind: what is in the box, a size, a real feature. Do not invent specifications. Check the final frame on a phone screen, since TikTok Shop's interface may cover part of the picture, and confirm its current safe areas on TikTok's pages.
Sources
Related posts
More in Media tools
- Caption a silent AI video: fixing caption_no_speech
A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.
- Caption a silent Seedance 2.5 clip with authored cues
A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.
- Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.
- Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.
Written by Sume