TikTok ad caption must match the video: check with a transcript
TikTok's ad policy requires the ad caption to be consistent with the video. How to transcribe an AI ad with Sume Video inspect and compare it with the caption.

TikTok's ad format policy has an Ad consistency section, and its second rule is that ad captions need to be consistent with the corresponding ad image or video. The page's examples of what is not allowed: the video says "Meet your future self!" while the caption says "Create a cartoon of yourself!", and the image says "Up to 30% off" while the caption says "Up to 50% off". With generated ads the caption and the video are often made in separate steps, so the mismatch is easy to create. Sume's Video inspect can transcribe the final cut so you can compare it with the caption text.
What does Ad consistency cover?
The section has three numbered rules. They are worth reading together, because the caption rule is the one that breaks in a generated workflow, but the first and third bite at the landing page.
| Rule | What the page says | Example it gives |
|---|---|---|
| 1 | Ad content, caption, text, images, videos and CTA must be consistent with the product or service on the landing page | Ad features product A, landing page shows product B |
| 2 | Ad captions must be consistent with the ad image or video | Video says "Meet your future self!", caption says "Create a cartoon of yourself!" |
| 3 | Display Name and App Name must match the promoted product, service or app on the landing page | Not given in the section's examples |
Why does a generated ad drift from its caption?
A caption is often written from a brief, while a video model follows the prompt it was given. If the brief says 50% off and the spoken line or burned text says 30%, nothing in a pipeline raises an error. The same is true when you re-cut: a trim or a Timeline pass changes what the video says, and the caption stays as it was.
The check is cheap compared with a rejected ad: transcribe the final file, read the spoken claims and any burned text, and compare numbers and promises with the caption line by line.
How do you transcribe the final cut with Sume?
Video inspect reads one clip on media.sume.com. Set transcribe: true and you get a transcript with text, words[] and, with segmentation.mode: "sentence", sentence segments[]. The public STT rate in the doc is $0.01 per audio minute. If you omit duration_seconds the hold reserves one minute; the hint maxes at 600 seconds. Every inspect is also billed for its Modal compute, up to its hold.
The call is synchronous by default: it waits up to 30 seconds and answers 200 with the finished inspect, or 202 with a job to poll. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-check-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/ad.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"segmentation": {"mode": "sentence"}
}'A short review routine
Run the check in this order. First, read the caption and write down every claim in it: product name, price, discount, deadline. Second, read the transcript and the stills and write down the same claims as the video makes them. Third, compare the two lists. A number that differs is a reject risk under rule 2; a product that differs from the landing page is a reject risk under rule 1.
Do it again after every re-cut. A trim of a few seconds can drop the sentence that carried the discount, and a new soundtrack changes nothing about the caption but can change what a viewer hears.
What about text burned into the picture?
A transcript only covers speech. Text burned into frames, such as a price badge, needs stills: ask Video inspect for frames at the moments the badge is on screen, or use Video frames at named times, then read them. Sume's docs describe both as returning image artifacts for you to look at; they do not read the text for you.
If you add the text yourself, Video captions burns authored cues or segments (each with text, start, end in seconds) without speech-to-text. A standalone caption job is $0.20 for videos up to 60 seconds per the doc. Keep the caption wording in one place and feed the same string to the cues and to the TikTok caption field, so the two cannot disagree.
Sources
Related posts
More in Use cases
- Editing a TikTok ad's video triggers a new review: plan for 24 hours
TikTok says saving or editing an ad's creative starts a new review, and most ads clear in about 24 hours. Plan a swapped AI video, and cut one with Video trim.
- Does a TikTok ad need sound? Find silent AI videos with Sume
TikTok's ad policy says an ad must contain audio that is not muffled. Many AI video clips are silent. Probe for audio with Sume and add a bed in Timeline.
- TikTok Ad Network carousel: 50 images max, pull stills from a video
TikTok Ad Network carousel and image ads take up to 50 images at 1200x628, 640x640 or 720x1280. Pull 24 stills per call with Sume video frames.
- TikTok Ad Network video: up to 10 minutes, 500 MB, 720x720 allowed
TikTok Ad Network (Pangle) video ads accept 1280x720, 720x1280 or 720x720, up to 10 minutes and 500 MB. Cut each size with Sume video trim.
Written by Sume