TikTok ad caption must match the video: check with a transcript

TikTok's ad policy requires the ad caption to be consistent with the video. How to transcribe an AI ad with Sume Video inspect and compare it with the caption.

4 min readSume
All posts

TikTok's ad format policy has an Ad consistency section, and its second rule is that ad captions need to be consistent with the corresponding ad image or video. The page's examples of what is not allowed: the video says "Meet your future self!" while the caption says "Create a cartoon of yourself!", and the image says "Up to 30% off" while the caption says "Up to 50% off". With generated ads the caption and the video are often made in separate steps, so the mismatch is easy to create. Sume's Video inspect can transcribe the final cut so you can compare it with the caption text.

What does Ad consistency cover?

The section has three numbered rules. They are worth reading together, because the caption rule is the one that breaks in a generated workflow, but the first and third bite at the landing page.

TikTok Ads Help, Ad format and functionality policy (Ad consistency), read 2026-10-03
RuleWhat the page saysExample it gives
1Ad content, caption, text, images, videos and CTA must be consistent with the product or service on the landing pageAd features product A, landing page shows product B
2Ad captions must be consistent with the ad image or videoVideo says "Meet your future self!", caption says "Create a cartoon of yourself!"
3Display Name and App Name must match the promoted product, service or app on the landing pageNot given in the section's examples

Why does a generated ad drift from its caption?

A caption is often written from a brief, while a video model follows the prompt it was given. If the brief says 50% off and the spoken line or burned text says 30%, nothing in a pipeline raises an error. The same is true when you re-cut: a trim or a Timeline pass changes what the video says, and the caption stays as it was.

The check is cheap compared with a rejected ad: transcribe the final file, read the spoken claims and any burned text, and compare numbers and promises with the caption line by line.

How do you transcribe the final cut with Sume?

Video inspect reads one clip on media.sume.com. Set transcribe: true and you get a transcript with text, words[] and, with segmentation.mode: "sentence", sentence segments[]. The public STT rate in the doc is $0.01 per audio minute. If you omit duration_seconds the hold reserves one minute; the hint maxes at 600 seconds. Every inspect is also billed for its Modal compute, up to its hold.

The call is synchronous by default: it waits up to 30 seconds and answers 200 with the finished inspect, or 202 with a job to poll. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-check-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/ad.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en",
    "segmentation": {"mode": "sentence"}
  }'

A short review routine

Run the check in this order. First, read the caption and write down every claim in it: product name, price, discount, deadline. Second, read the transcript and the stills and write down the same claims as the video makes them. Third, compare the two lists. A number that differs is a reject risk under rule 2; a product that differs from the landing page is a reject risk under rule 1.

Do it again after every re-cut. A trim of a few seconds can drop the sentence that carried the discount, and a new soundtrack changes nothing about the caption but can change what a viewer hears.

What about text burned into the picture?

A transcript only covers speech. Text burned into frames, such as a price badge, needs stills: ask Video inspect for frames at the moments the badge is on screen, or use Video frames at named times, then read them. Sume's docs describe both as returning image artifacts for you to look at; they do not read the text for you.

If you add the text yourself, Video captions burns authored cues or segments (each with text, start, end in seconds) without speech-to-text. A standalone caption job is $0.20 for videos up to 60 seconds per the doc. Keep the caption wording in one place and feed the same string to the cues and to the TikTok caption field, so the two cannot disagree.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume