YouTube's four disclosure triggers checked against Sume's tools
Four YouTube disclosure triggers: real person, altered real footage, invented realistic scene, music. Each is checked against Sume generation and media tools.

YouTube's page lists four things creators must disclose when the content looks realistic: making a real person appear to say or do something they did not do, altering footage of a real event or place, generating a realistic scene that did not occur, and creating music that is the main focus of the video. Most Sume generation falls under the third; trim and crop of your own footage touch none of them.
The list is from YouTube's disclosure page (read 2026-10-10). The mapping to Sume is my reading of what each Sume tool does, taken from its docs page.
The four triggers in YouTube's words
The page says creators must disclose AI-generated or meaningfully altered content that appears realistic and that makes a real person appear to say or do something they did not do, alters footage of a real event or place, generates a realistic scene that did not actually occur, or creates music that is the main focus of the video. The examples it gives for must-disclose are AI-generated footage of real locations, deepfakes of real people, and fabricated news events.
Each trigger against Sume
Sume has generation tools and media tools. The generation tools create content with a model. The media tools cut, crop, dim, assemble and caption, and their docs say they run only worker ffmpeg with no provider inference.
| YouTube trigger | Sume tool that can cause it | Disclose? |
|---|---|---|
| Real person appears to say or do something they did not | Image-to-video or avatar video using a real person's likeness | Yes, if realistic |
| Alters footage of a real event or place | Generation that restyles or extends real footage | Yes, if realistic |
| Realistic scene that did not occur | Text-to-video, image-to-video | Yes, if realistic |
| Music that is the main focus | Generated music tracks | Yes, if the music is the focus |
| None of the four | Video trim, crop, dim, Timeline of your own footage | No AI trigger from the tool |
Where Sume's honest limits sit
Two Sume points are worth stating plainly. Avatar 1.0 is English-only, and its subject should be a person who agreed to appear; a likeness of a real person saying words they did not say is exactly the first trigger. And voice cloning on Sume is available in the app, not through the API. YouTube's page separately says cloning your own voice for voice-overs or dubs does not need disclosure, so that case is narrow and personal.
Trim, crop and dim leave the content as filmed. Timeline puts clips in order and can add transitions and an audio spine. If every clip inside was filmed by you, the assembled video is not a synthetic one just because Timeline touched it. If one clip inside was generated, that clip can still trigger the third item.
A short checklist before upload
Ask four yes or no questions: does a real person appear to do something they did not, was real footage of a place or event altered, is there a realistic scene that never happened, and is the music the main point. Any yes on a realistic clip means disclose. Keep your Sume job kinds with the upload record; the result envelopes name them, for instance video_trim or timeline_render.
Notice that three of the four triggers hinge on the word realistic. The page's own counter-example is a fantasy scenario, which needs no disclosure. For a brand video, that means stylised illustration, obvious animation or an openly surreal look sits outside the rule, while a lifelike product demo generated from a text prompt sits inside it. The tool does not decide that; the look does.
The fourth trigger, music as the main focus, is the one people forget. A background bed under a talking-head clip is not the main focus, but a music video with a generated track is. If you build an audio spine in Timeline from generated music, ask whether a viewer would say the video is about the music.
- Generated and realistic: disclose.
- Generated and clearly fantastical: the page says no disclosure.
- Your own footage, only cut or cropped: no AI trigger from those tools.
Sources
Related posts
More in Comparisons
- YouTube's resolution ladder: which rows Sume Timeline can render
YouTube lists eight 16:9 sizes from 426x240 to 7680x4320. Sume Timeline renders even edges from 256 to 2160, so 1080p, 720p, 480p and 360p fit; the rest do not.
- YouTube: upscale and repair need no AI label, realistic edits do
YouTube's page exempts sharpening, upscaling or repair, beauty filters and colour changes from AI disclosure. Realistic synthetic people or places need it.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume