Label an AI cold open on screen with a two-second caption card
Metricool: 3 billion+ AI-labelled TikTok videos. Burn your own disclosure card onto a silent AI clip with /v1/video-captions cues, $0.20 for up to 60 seconds.

To put your own AI disclosure on screen for an AI-generated cold open, send the silent clip to POST /v1/video-captions with one authored cue, for example a card from 0 to 2 seconds that reads "AI-generated". That burns the text into a new video for $0.20 on a clip up to 60 seconds (read 2026-10-03), and it works on silent footage because cues skip speech-to-text.
The context is scale. Metricool's TikTok news page reports more than 3 billion AI-labelled videos on the platform (reported, read 2026-10-03). A series that opens on generated footage is in that group, and an on-screen card in your own words is a durable choice that travels with the file, whatever a platform does or does not add on upload. This article does not state any platform's rules; read the one you post to.
Why cues, and why silent clips
The caption route has four ways to supply text: script_text, words, cues and segments. They are mutually exclusive. cues (and segments) are phrase-level overlay cards with text, start and end in seconds, and they skip speech-to-text entirely, so Sume burns exactly the copy at exactly the times you give it.
That is the right tool for a model-generated cold open, which is usually silent or carries a sound bed rather than speech. A speech-based caption job on a silent clip fails with caption_no_speech, and the suggested next action is to use overlay captions. Cues are that overlay.
| Choice | Setting | Reason |
|---|---|---|
| Input field | cues | Skips speech-to-text, works on silent clips |
| Card text | Short, plain wording | Readable in two seconds on a phone |
| Timing | start 0, end 2 | Visible from the first frame |
| Style | A bold Latin style such as slam | Latin copy only; Hangul needs a Hangul style |
| Price | $0.20 up to 60 s | Fixed estimate for the standalone caption job |
The request
The clip must be a public HTTPS video URL. Submit with an Idempotency-Key, then read the job by id. The route is job-backed, so a successful submit gives you an id to poll rather than a finished file.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {"video_url": os.environ["COLD_OPEN_URL"],
"style": "slam",
"cues": [{"text": "AI-generated", "start": 0.0, "end": 2.0}]}
r = requests.post(f"{API}/v1/video-captions",
headers={**H, "Idempotency-Key": "ep1-label-v1"}, json=body)
r.raise_for_status()
rid = r.json()["request_id"]
for _ in range(60):
s = requests.get(f"{API}/v1/jobs/{rid}/status", headers=H).json()
if s.get("status") in ("completed", "failed", "canceled"):
break
time.sleep(3)
print(s.get("status"))
print(requests.get(f"{API}/v1/jobs/{rid}/result", headers=H).json())What to do in a series
Label the opening, not the whole episode, unless you have a reason to. A two-second card on the first clip is enough for a viewer to see it, and the cost is a single caption job per episode. For eight episodes that is $1.60.
Do the burn before you cut the clip into Timeline. The card then belongs to the cold open, and Timeline treats the captioned file like any other clip. If you label after the join, the $0.20 figure covers videos up to 60 seconds, so check the live catalog for the price of a longer cut.
Do not mix this with speech captions on the same clip. One job takes one text field. If the episode is narrated and you want both words and a label, caption the narration with script_text and add the label to the silent cold open, which is a separate clip. That keeps each job simple and each result checkable.
Keep the wording and the job id in your episode record. If you change the wording later, run a new job with a new key such as ep1-label-v2 on the original clean clip, not on the already-captioned file, so you do not stack two cards.
A worked example: a season of eight episodes, each opening on a 6-second silent cold open. Eight caption jobs at $0.20 is $1.60, and each takes the clean clip as its input, so a wording change for the whole season is another $1.60 and no new video generation. Compare that with regenerating eight cold opens, which costs far more than a text overlay and risks changing the footage you had approved.
Last, look at the output once on a phone. A card that reads well on a desktop preview can sit under the platform's own interface in a vertical feed; move the placement with the design.placement.anchor_ratio override if it does.
What Sume does and does not do
Sume burns the card you write into a new video and records the job. It does not decide whether a clip needs a disclosure, it does not add or verify any platform's AI label, and it does not make a legal determination. The wording and the decision to label are yours.
Sources
Related posts
More in Use cases
- Lantern Festival 2027 riddle video: five riddles, one 35-second cut
The Lantern Festival is 20 February 2027. Make a lantern-riddle video from five stills, one motion clip and caption cues for about $1.22 in Sume costs.
- Las Posadas December 16-24: nine nightly invite videos for $1.93
Las Posadas runs nine nights from December 16. Reuse one invite still and burn nine different host lines with nine caption jobs for about $1.93 on Sume.
- Legal 77.3%, financial 78.4%: a review step before an avatar render
HeyGen's survey puts avatar use at 77.3% in legal and 78.4% in financial services. Add a script sign-off and job record to a Sume avatar workflow before render.
- LLM-written cue sheet: validate the times before you burn captions
Gemini 3.8 Flash or any LLM can draft caption cues as JSON. Check order, overlap and length in Python, then burn them with /v1/video-captions for $0.20.
Written by Sume