Does a lip-synced dub in your own voice need YouTube's AI label?

YouTube says cloning your own voice for voice-overs or dubs needs no disclosure; a real person made to say what they never said does. How Sume lip sync fits.

3 min readSume
All posts

YouTube's help page on disclosing altered or synthetic content lists "cloning one's own voice to create voice overs or dubs" as something creators do not need to disclose. The same page requires disclosure when AI makes a real person appear to say or do something they did not do. A lip-synced dub in your own voice, on your own face, sits on the first side of that line; a lip-synced line put in someone else's mouth sits on the second.

What the page actually says

The page (read 2026-10-09) says creators must disclose realistic altered or synthetic content, giving as examples making a real person appear to say or do something they did not do and altering footage of a real event or place. It exempts non-realistic content and minor edits, such as beauty filters, color adjustment and lighting filters.

It does not mention lip sync by name, and it does not say whether re-animating a mouth to match a dub counts as a minor edit. So the page supports the voice-dub case and leaves the visual case to your judgment against the 'real person appears to say something they did not' test.

It helps to separate three questions. Is the voice yours? Is the face a real person's? Do the words match what that person actually said or would say? YouTube's page keys on the second and third: a realistic depiction of a real person saying something they did not say. A dub of your own words in your own voice does not change what you said, only the language or the take.

YouTube's wording applied to common lip-sync cases (YouTube Help, read 2026-10-09)
CaseWhat the page saysReading
Your own cloned voice dubbing your own videoListed as not requiring disclosureNo label needed per the page
A real person shown saying words they never saidRequires disclosureLabel it
A synthetic face that does not depict a real personJudgment call: does a real person appear to say something they did not?Decide and record your reasoning

Where Sume's lip-sync routes come in

Sume lists MiniMax H3 Max Lip Sync at $0.10 per audio second at 768p and VEED Fabric 1.0 at $0.1875 per audio second at 720p. Both take a still and a Sume-hosted audio file. The audio is yours to supply; the routes do not decide what your clip says or whose voice it is.

H3 Max needs 5 to 14.8 seconds of audio, so a 12-second dub at 768p costs 12 x $0.10 = $1.20.

Keep a short record per video: the audio source, whose voice and face are used, the Sume job id, and the date. That record answers a platform query in minutes and costs nothing. It also makes a later change of YouTube's wording easy to check against what you did.

This page is a reading of YouTube's rule, not legal advice, and YouTube can change the wording. Platform rules are also separate from local law: some jurisdictions have their own rules on synthetic media and likeness that apply regardless of what a platform asks. Sume's API documentation for the avatar and lip-sync routes covers inputs, limits and prices, and does not set a consent policy, so the decision about what you may make and publish stays with you and the people involved.

  • Keep the source consent and the audio origin with the job id.
  • If a real person speaks words they did not say, use YouTube's disclosure setting.
  • Re-read the page before a campaign; it is YouTube's, not Sume's, to change.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume