Label an AI cold open on screen with a two-second caption card

Metricool: 3 billion+ AI-labelled TikTok videos. Burn your own disclosure card onto a silent AI clip with /v1/video-captions cues, $0.20 for up to 60 seconds.

5 min readSume
All posts

To put your own AI disclosure on screen for an AI-generated cold open, send the silent clip to POST /v1/video-captions with one authored cue, for example a card from 0 to 2 seconds that reads "AI-generated". That burns the text into a new video for $0.20 on a clip up to 60 seconds (read 2026-10-03), and it works on silent footage because cues skip speech-to-text.

The context is scale. Metricool's TikTok news page reports more than 3 billion AI-labelled videos on the platform (reported, read 2026-10-03). A series that opens on generated footage is in that group, and an on-screen card in your own words is a durable choice that travels with the file, whatever a platform does or does not add on upload. This article does not state any platform's rules; read the one you post to.

Why cues, and why silent clips

The caption route has four ways to supply text: script_text, words, cues and segments. They are mutually exclusive. cues (and segments) are phrase-level overlay cards with text, start and end in seconds, and they skip speech-to-text entirely, so Sume burns exactly the copy at exactly the times you give it.

That is the right tool for a model-generated cold open, which is usually silent or carries a sound bed rather than speech. A speech-based caption job on a silent clip fails with caption_no_speech, and the suggested next action is to use overlay captions. Cues are that overlay.

A disclosure card on a silent cold open (read 2026-10-03)
ChoiceSettingReason
Input fieldcuesSkips speech-to-text, works on silent clips
Card textShort, plain wordingReadable in two seconds on a phone
Timingstart 0, end 2Visible from the first frame
StyleA bold Latin style such as slamLatin copy only; Hangul needs a Hangul style
Price$0.20 up to 60 sFixed estimate for the standalone caption job

The request

The clip must be a public HTTPS video URL. Submit with an Idempotency-Key, then read the job by id. The route is job-backed, so a successful submit gives you an id to poll rather than a finished file.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

body = {"video_url": os.environ["COLD_OPEN_URL"],
        "style": "slam",
        "cues": [{"text": "AI-generated", "start": 0.0, "end": 2.0}]}
r = requests.post(f"{API}/v1/video-captions",
                  headers={**H, "Idempotency-Key": "ep1-label-v1"}, json=body)
r.raise_for_status()
rid = r.json()["request_id"]

for _ in range(60):
    s = requests.get(f"{API}/v1/jobs/{rid}/status", headers=H).json()
    if s.get("status") in ("completed", "failed", "canceled"):
        break
    time.sleep(3)
print(s.get("status"))
print(requests.get(f"{API}/v1/jobs/{rid}/result", headers=H).json())

What to do in a series

Label the opening, not the whole episode, unless you have a reason to. A two-second card on the first clip is enough for a viewer to see it, and the cost is a single caption job per episode. For eight episodes that is $1.60.

Do the burn before you cut the clip into Timeline. The card then belongs to the cold open, and Timeline treats the captioned file like any other clip. If you label after the join, the $0.20 figure covers videos up to 60 seconds, so check the live catalog for the price of a longer cut.

Do not mix this with speech captions on the same clip. One job takes one text field. If the episode is narrated and you want both words and a label, caption the narration with script_text and add the label to the silent cold open, which is a separate clip. That keeps each job simple and each result checkable.

Keep the wording and the job id in your episode record. If you change the wording later, run a new job with a new key such as ep1-label-v2 on the original clean clip, not on the already-captioned file, so you do not stack two cards.

A worked example: a season of eight episodes, each opening on a 6-second silent cold open. Eight caption jobs at $0.20 is $1.60, and each takes the clean clip as its input, so a wording change for the whole season is another $1.60 and no new video generation. Compare that with regenerating eight cold opens, which costs far more than a text overlay and risks changing the footage you had approved.

Last, look at the output once on a phone. A card that reads well on a desktop preview can sit under the platform's own interface in a vertical feed; move the placement with the design.placement.anchor_ratio override if it does.

What Sume does and does not do

Sume burns the card you write into a new video and records the job. It does not decide whether a clip needs a disclosure, it does not add or verify any platform's AI label, and it does not make a legal determination. The wording and the decision to label are yours.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume