YouTube AI label: player or description, by how real it looks
YouTube shows the altered or synthetic label in the player for photorealistic content, in the expanded description for animated. What it means.

On YouTube, photorealistic altered or synthetic content gets the disclosure label in the video player, and non-photorealistic or animated content gets it in the expanded description. The visible spot depends on how realistic the clip looks, so decide that when you plan the clip, not after upload.
Where the label appears
YouTube's help page on disclosing altered or synthetic content separates the two cases. The table summarizes what it says, read on 2026-10-03.
| Content | Where the label shows |
|---|---|
| Photorealistic altered or synthetic content | In the video player |
| Non-photorealistic or animated content | In the expanded description |
| Content YouTube labels itself | YouTube may apply a label for its own GenAI tools, C2PA metadata or internal detection |
What still needs disclosure at all
The same page lists what requires the toggle: a real person saying or doing something they did not, an altered real event or place, a realistic scene that did not occur, and music as the main focus. It does not require disclosure for non-realistic content, beauty filters, color or lighting adjustments, captions, or script help.
Choosing a look on purpose
A team that generates many clips can use that split as a creative choice. A stylized or illustrated look moves the label out of the player, while a photorealistic look keeps it in the player. Neither choice removes the duty to disclose when the content falls in the listed categories. Per the page, disclosure itself does not limit reach or monetization, and repeated non-disclosure can lead to removal or suspension from the Partner Program.
Do not pick a look to dodge a label. Pick the look the story needs, then disclose according to the table.
Pin the look in the request
Sume's video API takes a text prompt, an aspect_ratio and optional frame_images or input_references. A style written into the prompt and a stylized first frame both push the output the same direction across a batch, which keeps a series consistent in whether it reads as photographic or animated.
Keep a record of the planned disclosure state per clip in your own tracking sheet.
{
"model": "sume/auto",
"prompt": "Flat illustrated explainer scene, hand-drawn look, a paper boat crossing a river",
"aspect_ratio": "9:16",
"duration": 5
}Sources
Related posts
More in Use cases
- YouTube end screens need a 25-second video: plan the clip length
YouTube end screens need videos 25 seconds or longer. Most Sume video models stop at 15 seconds, so here is how to reach 25 with long models or Timeline.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
Written by Sume