Swipe up text in a TikTok ad: why it breaks policy, what to use

TikTok's ad policy names 'Swipe up to learn more' as an unsupported gesture. Keep fake gestures out of AI ad video and burn a real call to action with Sume.

4 min readSume
All posts

Do not put "Swipe up to learn more" in a TikTok ad. The ad policy lists invalid buttons, induced gestures or text that portray unsupported functionality, and names that text as its example, because swiping up just leads users to the next video and not anywhere else. The risk with AI video is that models trained on social clips draw swipe arrows, fake buttons and "link in bio" style overlays on their own. Check the frames, and add your call to action yourself with Sume Video captions.

What does the policy say about fake gestures?

The rule sits under Ad and editorial quality, alongside the visual-clarity rules. The wording groups three things: buttons, gestures and text. The page's own example is the swipe-up line. It does not list every phrase, so use the principle: if the video tells a viewer to do something the app does not do, it is an induced gesture.

TikTok Ads Help, Ad format and functionality policy (Ad and editorial quality), read 2026-10-03
Element in the videoWhy the policy objects
"Swipe up to learn more." textSwiping up leads to the next video, not to a destination (the page's own example)
A drawn button that does nothingPage lists invalid buttons that portray unsupported functionality
A hand-gesture promptPage lists induced gestures
Blurred or masked third-party watermarkListed separately as not allowed

Where do fake gestures come from in AI video?

A generated clip can include arrows, finger icons or a drawn interface bar that nobody asked for, and because the picture is generated you cannot rely on the prompt alone to keep them out. TikTok's Reach and Frequency reservation page separately says a non-Spark ad must not mimic the TikTok interface, so a drawn app chrome is a second risk.

Put the negative guidance in your prompt if the model supports it, but treat the rendered file as the thing to check.

How do you check and replace with Sume?

Run Video frames at 1 fps (fps is above 0 up to 2, 24 frames per call, source up to 300 seconds) and look at the lower and upper parts of every still. If a gesture prompt appears, re-render the shot or trim it out with Video trim.

Then add the call to action as authored text. Video captions takes cues with text, start and end, skips speech-to-text and burns them in the style you name. Use wording that matches what the ad's button and landing page really do, for example "Shop now" next to a shop CTA.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cta-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "slam",
    "cues": [{"text": "Shop now", "start": 6.0, "end": 9.0}]
  }'

What should the ending of the ad do instead?

The policy's objection is to a gesture that does nothing, not to a call to action. The ad format has its own supported call-to-action button, and the video can simply say what the offer is. A clear last two to three seconds with the product, the offer and a plain verb is enough, and it keeps the caption, the video and the landing page telling one story.

If you build the ad in Timeline, give the end card its own slot, with a fit of cover so the end card fills the 1080x1920 default output, and burn the line afterwards. That separates the picture, which you check for stray gestures, from the wording, which you write.

Does the CTA text also have to be consistent?

Yes. The same policy's Ad consistency section says the ad's text, video and call to action need to be consistent with the promoted product on the landing page, and that the caption must match the video. So the burned line, the caption and the platform CTA button should say the same thing.

A short habit helps: keep the CTA line in one variable, use it in the cue and in the caption field, and review the final video's last frames on a phone. The caption job is $0.20 for a clip up to 60 seconds, so a fix after review costs less than a rejected flight.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume