Put a promo code on screen in an AI product video with caption cues
Burn a Black Friday code into a silent clip with $0.20 caption cues, key the retry by SKU and code, then pull stills with video-frames to check the text.

To put a promo code on screen in an AI-generated product video, do not ask the video model to draw the letters. Generated clips garble small text, and a wrong code costs you a sale. Instead, render the clip silent, then burn the code with the Sume video captions API using authored cues, which place text you wrote at times you choose, with no speech recognition involved. The job costs $0.20 for a video of 60 seconds or less (read 2026-10-05).
Why cues, not a transcript
A silent image-to-video clip has no speech, so a plain caption job fails with caption_no_speech. The docs say to send cues (or segments) with text, start and end to burn authored overlay text without ASR. Because the text is yours, the code string is exact every time.
Meta says Reels default to sound on, so a spoken code is also reasonable, but Meta's Reels ads page pairs 9:16 video with key messages inside the safe zone. On-screen text covers viewers who scroll with sound off. Set design.placement.anchor_ratio to move the line away from platform buttons.
A runnable helper
The helper below posts one caption job per SKU. The Idempotency-Key combines SKU and code, so retrying the same pair replays the first job instead of billing again, while a new code on the same SKU is a new job.
import json, os, urllib.request
KEY = os.environ["SUME_API_KEY"]
def burn_code(video_url, sku, code):
body = {"video_url": video_url, "style": "punch", "cues": [
{"text": "BLACK FRIDAY", "start": 0.0, "end": 2.0},
{"text": "Code " + code + " at checkout", "start": 2.0, "end": 4.5},
]}
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + KEY,
"Content-Type": "application/json",
"Idempotency-Key": "promo-" + sku + "-" + code})
with urllib.request.urlopen(req) as resp:
return json.load(resp)
if __name__ == "__main__":
url = "https://media.sume.com/artifacts/demo/clip.mp4"
print(burn_code(url, "sku1", "BF26"))
Check the pixels, not the request
After the job finishes, extract two stills with video frames: POST /v1/video-frames with at: [3.0], one frame per cue window. Look at the stills, or run your own text check against the known code. Video frames is billed by Modal compute, not a fixed price, so treat it as a small line and read the actual usage afterwards.
| Check | Source of the rule | What to confirm |
|---|---|---|
| Code spelling | Your cue text | Exact string, same case |
| Cue timing | Caption job cues | Code visible 2 s or more |
| Non-Spark TikTok ads | TikTok Ads Help | Ad captions (the ad text) are white, fixed font, no links, @ or hashtags |
| Spark Ads caption | TikTok Ads Help | Ad caption up to 4 lines |
Keep the ad text separate from the burned-in cues: the TikTok ad specifications say the ad captions of non-Spark ads are white in a fixed font and carry no clickable links, @ or hashtags, so put the code in the video and not only in the ad text. At $0.20 a job, 100 SKUs cost $20.00 for one code; changing the code means paying the caption line again, so decide the code before the batch, as the 100-ad finishing budget shows.
Sources
Related posts
More in Media tools
- Proof frame for a 1080x1920 export: video frames PNG at 1920
Pull a lossless 1080x1920 still from your vertical export with video frames (format png, max_edge 1920) to check captions and crop before you upload.
- Silent reference clip: reference_ingest's -60 LUFS gate and STT
reference_ingest marks a clip audio.silent at -60 LUFS or below and skips STT, so a silent reference costs no transcript. What to do next: plan new music.
- reference-ingest source_no_video_stream: audio-only file as input
source_no_video_stream means the reference file has no video track, such as an m4a. Send a clip with picture, or use audio detach to work with the sound.
- reference_ingest_stt_required: speech.language_code without STT
speech.language_code and duration_seconds only apply when the read transcribes. Set speech.allow_billed_stt true, use reference_remix, or drop them.
Written by Sume