YouTube: upscale and repair need no AI label, realistic edits do
YouTube's page exempts sharpening, upscaling or repair, beauty filters and colour changes from AI disclosure. Realistic synthetic people or places need it.

YouTube's disclosure page says video sharpening, upscaling or repair, voice or audio repair, beauty filters and colour adjustments do not need disclosure, while a realistic synthetic person, place or event does. Sume's trim, crop, dim and Timeline fall in neither list by name, so treat them as plain edits on your own footage.
These items are quoted from YouTube's page (read 2026-10-10). I am not adding classifications that the page does not make.
The not-required list on YouTube's page
The page lists non-realistic content and minor edits that do not need disclosure: fantasy scenarios, aesthetic-only changes such as beauty filters and colour adjustments, cloning your own voice to create voice-overs or dubs, production assistance such as script or thumbnail generation, and video sharpening, upscaling or repair and voice or audio repair. Its examples of no disclosure needed include gameplay footage, green screen effects and caption creation.
| On YouTube's no-disclosure list | Closest Sume tool | Note |
|---|---|---|
| Colour adjustments | Video filter dim | Dim only scales brightness; other tone filters need the filtergraph |
| Caption creation | Video captions | Burns text onto the clip |
| Video sharpening, upscaling or repair | Not a trim, crop or Timeline function | Not part of trim, crop or Timeline; check the live catalog for any other tool |
| Script or thumbnail generation | Image and text tools | Production assistance |
| Cloning your own voice for voice-overs | Voice cloning, app only | Not available through the API |
Trim and crop are not on either list
YouTube's page does not mention trimming, cropping or reframing. They are cuts and framing changes, not synthetic content, and the page's logic is realism plus alteration of what happened. A trim leaves every frame as shot. A crop from 16:9 to 9:16 shows a part of the same frame. Neither generates anything.
That is still my reading, not YouTube's wording. If your edit makes a real person appear to say something they did not, for example by cutting a sentence out of context, the first disclosure trigger is about that effect, whatever tool you used.
Where realistic edits start
The line on the page is realism. A tool that restyles a scene so it looks like a real event that never occurred sits on the disclose side. A generated stand-in for a real place does too. A filter that only changes colour does not. Sume's generation endpoints are on the first side whenever the output looks real; its media tools are on the second side because they only reorganise or darken the pixels.
Practical handling
Keep two lists per project: footage that came from a camera and clips that came from a model. Only the second list drives the disclosure answer. When both are in one Timeline, the answer follows the generated parts. Remember that the consequence for consistently skipping a needed disclosure is a manual label or penalties, including removal or Partner Program suspension, so a simple yes is cheaper than a dispute.
There is a useful test hiding in the page's examples. Every item on the no-disclosure list changes how the footage looks or sounds without changing what the footage claims to show. Every item on the must-disclose list changes what the footage claims to show: a person saying something, a place that looks different, a scene that never happened. Trim, crop and dim sit on the first side, while generation sits on the second.
Mixed projects are common. A creator might shoot real footage, dim it, trim it, drop in one generated B-roll shot, and assemble everything in Timeline. The answer to the disclosure question is then driven by the B-roll shot, not by the five other steps. Writing that down once per project is easier than rethinking it for every upload.
Sources
Related posts
More in Comparisons
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
- HeyGen alternatives with an API: price units, limits, and fit
HeyGen alternatives with an API: Synthesia, Creatify, Argil, Arcads, and Sume compared by price unit, API shape, limits, and live vs rendered avatars.
- Sume vs Higgsfield: two video agents compared on API, MCP, and price
Higgsfield and Sume both put a creative agent over many video models. How their agents, APIs, MCP servers, plans, and billing differ, from each vendor's pages.
Written by Sume