Sound-off holiday video: burn your own text cues with Sume captions
For a silent product clip, send cues with text, start and end to POST /v1/video-captions. No speech-to-text runs. Code sample, $0.20 per job, and the limits.

To put text on a silent holiday clip, call POST /v1/video-captions with cues, each having text, start and end in seconds. The Sume docs say that with cues, segments, words or script_text forms you control the text, and that cues run without speech-to-text, so they work on clips with no audio. Without them, a silent clip fails with caption_no_speech. Public price is $0.20 per job for videos up to 60 seconds.
Pick the right text input
Send only one of these per request.
| Field | What it does | Needs speech? |
|---|---|---|
| script_text | Aligns your script to STT word timings | Yes |
| words | Word-level text you time yourself | No |
| cues / segments | Phrase-level overlay cards with text, start, end | No |
| (none) | STT transcribes the audio | Yes |
Steps
The captions endpoint takes a public HTTPS video URL. Put the clip URL in CLIP_URL and your key in SUME_API_KEY. Cue times below are examples, so set yours from the length of your clip.
import json, os, urllib.request
body = {
"video_url": os.environ["CLIP_URL"],
"style": "black-outline",
"cues": [
{"text": "Free gift wrap", "start": 0.5, "end": 2.5},
{"text": "Ships by Dec 20", "start": 2.5, "end": 5.0},
],
}
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": "sound-off-cues-001"},
)
print(json.load(urllib.request.urlopen(req)))The response is a job. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result, as with the other jobs. Latin text defaults to the slam style.
Placement
The design.placement.anchor_ratio field sets where the line sits as a fraction of frame height, so you can move the text clear of any platform interface. The punch and tiktok-green styles do not support design.
Writing cues that read in a glance
A cue is a card of text that appears from start to end. Keep each to a few words, hold it long enough to read, and make sure the cues do not overlap. A common rule of thumb is about a second for the first three words and a little more for each extra word, but treat that as a starting point and watch the result.
For a holiday clip, a three-cue structure works: the offer, the proof point, and the date or call to action. For example, "Free gift wrap", "Ships by Dec 20" and "Shop now". Put the cue text in your own copy deck so the team can approve it before it is burned in.
If you decide on a new line after the render, use source_caption_id with new cues or words. The docs say a restyle reuses the source video, so you do not upload the clip again.
Budget and detail
One more point on cost: because cues skip speech-to-text, the job is the same $0.20 for a clip up to 60 seconds as a spoken one. If you have 50 silent SKU clips, that is $10.00 for the captions, and if you also render them through Timeline first it is another $0.10 each, so $15.00 in total for the batch. Write all the cue copy before you start, so you do not pay to re-burn a typo.
What Sume does not do
Sume does not mark where platform interface elements sit, and it does not pick cue timing for you. It also does not make a separate caption file.
Sources
Related posts
More in Use cases
- 45,000-character script: three Sume TTS jobs plus a join for $2.17
A 45,000-character script needs three TTS requests under the 20,000 cap. At 72 cents each plus a one-cent join the audio costs $2.17 on Sume.
- Sports club signup ad: Kling 3, 15 s, 1080p, $3.15 with sound
A 15-second 1080p signup ad on Kling 3 costs $2.10 silent or $3.15 with sound on Sume, the same as at 720p. Where it fits against Wan 3.0 and Seedance 2.5.
- Square YouTube Short from vertical AI clips: 1080x1080 on Timeline
Render a square Short from 9:16 AI clips by setting output width and height on Timeline 1.0, choosing a fit mode, with the cost for 60, 120 and 180 seconds.
- Swap a product background with qwen-image: 2.5 cents, one reference
qwen-image costs $0.025 per image on Sume and takes up to 10 references. A background-swap request, the 1,000-photo bill, and what qwen-image-max cannot do.
Written by Sume