Face-swap clip on YouTube: label needed? Sume's beta caps at 15 s
YouTube asks for a label when a real person appears to say or do something they did not. Sume's face swap beta takes about 4 to 15 seconds; $3.68 max at plus.

Short answer
A face swap that puts a real person's face on someone else's footage is the case YouTube's help page leads with: content that makes a real person appear to say or do something they did not do. Plan to disclose it. Sume's Avatar Face Swap 1.0 is a Beta endpoint that applies a ready avatar face to a public source video of roughly 4 to 15 seconds with usable audio, and its largest cost estimate at plus is $3.68.
What the YouTube page says
The page lists realistic synthetic content that needs a label and that first trigger is the one a face swap onto someone else's footage most plainly meets, in our reading. It lists cosmetic edits, script generation and upscaling as exempt. It also says disclosure does not limit reach or monetization eligibility, and that creators who consistently do not disclose may face a manual label or penalties.
| Point | Page statement |
|---|---|
| Real person appears to say or do something | Disclose |
| Cosmetic edits, upscaling | No disclosure needed |
| Disclosing | Does not limit reach or monetization eligibility |
| Repeated non-disclosure | Manual label or penalties, including Partner Program suspension |
What Sume's Beta takes
The endpoint is POST /v1/models/sume/avatar-face-swap/v1.0/runs. The required fields are avatar_handle, video_url and quality, and quality has no default in Beta. The video_url must be a fetchable public HTTPS video; signed or private URLs and provider task URLs are rejected. It does not accept prompts, transcripts, duration knobs or aspect ratio.
| Constraint | Value |
|---|---|
| Source video length | About 4 to 15 seconds |
| Audio | Usable audio needed |
| quality | Required: standard, plus or max |
| Result | video_url under media.sume.com |
Cost at the 15-second ceiling
Sume prices the Beta against the Avatar Video no-product rates. At the 15-second maximum, standard is 15 x $0.184 = $2.76, plus is 15 x $0.245 = $3.68, and max is 15 x $0.55 = $8.25. A shorter clip costs proportionally less. These are estimates for the largest allowed source video.
| Tier | Rate per second | 15 seconds |
|---|---|---|
| standard | $0.184 | $2.76 |
| plus | $0.245 | $3.68 |
| max | $0.55 | $8.25 |
Practical rules
Use only avatars you have the right to use, and only source footage you have the right to edit. If the avatar is a real person, get written agreement for this use. Then set the YouTube label when you upload, and add a disclosure in the clip or its description. This post is not legal advice, but the platform page is clear enough on the face-replacement case that a label is the safe default.
If you want a look without a real person's face, there is another route. Create an avatar from a prompt or profile and generate a script-driven video with Avatar Video. That does not take your footage, so you control the words and there is no original performer whose face is replaced. The trade is that you get a new scene rather than your existing clip.
Sources
Related posts
More in Sume Avatar 1.0
- Face swap or avatar video? Pick by whether you already have footage
Sume's face swap Beta applies an avatar face to your own 4-15 second video; Avatar Video renders a new clip from a script, up to 60 seconds. How to choose.
- Five 20-word avatar sentences make five 8-second clips, not three
Why five 20-word sentences plan as five 8-second clips (40 s) in Avatar 1.0, and how pairing shorter sentences saves clips.
- One wrong sentence in an avatar script: what 5 re-renders cost
Changing a script means a new Avatar 1.0 job. Five full re-renders of a 20-second take cost $18.40 on standard, $24.50 on plus and $55.00 on max.
- Griffin-Lite 48% is 26 of 54 callers: how wide is the range?
Tavus reports 48% of 54 callers took Griffin-Lite for human. A 95% Wilson interval on 26 of 54 is about 35% to 61%. What that means for buyers.
Written by Sume