Sound-off holiday video: burn your own text cues with Sume captions

For a silent product clip, send cues with text, start and end to POST /v1/video-captions. No speech-to-text runs. Code sample, $0.20 per job, and the limits.

3 min readSume
All posts

To put text on a silent holiday clip, call POST /v1/video-captions with cues, each having text, start and end in seconds. The Sume docs say that with cues, segments, words or script_text forms you control the text, and that cues run without speech-to-text, so they work on clips with no audio. Without them, a silent clip fails with caption_no_speech. Public price is $0.20 per job for videos up to 60 seconds.

Pick the right text input

Send only one of these per request.

Caption inputs on Sume, per the docs read 2026-10-08
FieldWhat it doesNeeds speech?
script_textAligns your script to STT word timingsYes
wordsWord-level text you time yourselfNo
cues / segmentsPhrase-level overlay cards with text, start, endNo
(none)STT transcribes the audioYes

Steps

The captions endpoint takes a public HTTPS video URL. Put the clip URL in CLIP_URL and your key in SUME_API_KEY. Cue times below are examples, so set yours from the length of your clip.

import json, os, urllib.request

body = {
    "video_url": os.environ["CLIP_URL"],
    "style": "black-outline",
    "cues": [
        {"text": "Free gift wrap", "start": 0.5, "end": 2.5},
        {"text": "Ships by Dec 20", "start": 2.5, "end": 5.0},
    ],
}
req = urllib.request.Request(
    "https://api.sume.com/v1/video-captions",
    data=json.dumps(body).encode(),
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Content-Type": "application/json",
             "Idempotency-Key": "sound-off-cues-001"},
)
print(json.load(urllib.request.urlopen(req)))

The response is a job. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result, as with the other jobs. Latin text defaults to the slam style.

Placement

The design.placement.anchor_ratio field sets where the line sits as a fraction of frame height, so you can move the text clear of any platform interface. The punch and tiktok-green styles do not support design.

Writing cues that read in a glance

A cue is a card of text that appears from start to end. Keep each to a few words, hold it long enough to read, and make sure the cues do not overlap. A common rule of thumb is about a second for the first three words and a little more for each extra word, but treat that as a starting point and watch the result.

For a holiday clip, a three-cue structure works: the offer, the proof point, and the date or call to action. For example, "Free gift wrap", "Ships by Dec 20" and "Shop now". Put the cue text in your own copy deck so the team can approve it before it is burned in.

If you decide on a new line after the render, use source_caption_id with new cues or words. The docs say a restyle reuses the source video, so you do not upload the clip again.

Budget and detail

One more point on cost: because cues skip speech-to-text, the job is the same $0.20 for a clip up to 60 seconds as a spoken one. If you have 50 silent SKU clips, that is $10.00 for the captions, and if you also render them through Timeline first it is another $0.10 each, so $15.00 in total for the batch. Write all the cue copy before you start, so you do not pay to re-burn a typo.

What Sume does not do

Sume does not mark where platform interface elements sit, and it does not pick cue timing for you. It also does not make a separate caption file.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume