YouTube AI disclosure for AI music and voiceover: what to keep

YouTube asks creators to disclose meaningfully altered or synthetic realistic content. What to keep from Sume music and TTS jobs, and what Sume cannot decide.

5 min readSume
All posts

Whether you must tick YouTube's altered-or-synthetic box is YouTube's decision, not Sume's, and the useful thing Sume can give you is a record of what you generated. YouTube's help page says creators must disclose when they use AI to meaningfully alter or generate photorealistic content; it describes some uses, such as caption creation and audio enhancement, as not needing disclosure.

This page does not interpret that policy for your channel. It lists what to keep from Sume Music and TTS jobs so you can answer the question quickly, and cites the YouTube Help page (read 2026-10-10). Re-read the page on the day you publish: platform rules change.

What does the YouTube page say?

Only the lines this post relies on are listed. The page covers more cases than a blog table can hold, including examples that do and do not require a label, so open it for the full text.

YouTube Help, as summarised (read 2026-10-10)
TopicWhat the page says
Meaningfully altered or generated photorealistic contentCreators are required to disclose it
Cloning your own voice for a voiceoverDescribed as not requiring disclosure
Caption creation and audio enhancementDescribed as not requiring disclosure

What can Sume tell you about a track or a voice?

Sume can show that a file came from a job: the prompt, the model route, the date and the artifact. For music, Music 1.0 returns an audio/mpeg artifact from Lyria 3.5. For speech, a TTS job returns audio and, if requested, word timings. Neither tells YouTube anything on its own; the label is set when you upload.

  • Music: prompt, date, job.request.routed_model if you used the router, the artifact.
  • Speech: transcript, voice id or avatar handle, language, engine id if you used the router.
  • Both: the plan you were on, since the Terms tie commercial use to paid plans.

A keep-list per video

Make one folder or database row per video. Put the final upload date, the platform setting you chose (label on or off) and the reason in one line. Then attach the generation records. If a viewer or the platform asks later, you can show what was machine-made and what was recorded by a person.

Suggested record per video (Sume docs for the fields, checked 2026-10-10)
ItemSource
Music prompt and artifactMusic 1.0 job
Engine that ranjob.request.routed_model (router jobs)
Voice and transcriptTTS job request
Caption jobVideo captions job
Label decision and reasonYour note

What stays outside Sume?

Sume's docs do not state whether any given video needs a YouTube label, whether a track will match Content ID, or how another platform treats AI audio. Treat those as open and check the platform page each time. A stored page compares label rules across TikTok, Instagram, YouTube and Meta with a grade for each source, which is a faster start than reading five help centres.

A worked example

Say you publish a 45 second explainer with a Sume Music bed and a TTS voiceover of your own script. Your record holds one music prompt, one transcript, one voice id, two job results and the upload settings. That is eight lines of text, and it takes five minutes to fill in at upload time. Skipping it is what makes a later question expensive: without the prompt and date you cannot show when or how the audio was made.

If you cloned your own voice in the Sume app and use it for the voiceover, note that too. YouTube's page describes cloning your own voice as a case that does not need disclosure, but the same page is the one to re-read when your use is different, for example a voice that is not yours.

  • Fill the record at upload time, not months later.
  • Store it with the video file, not in chat history.
  • Update it if you swap the music or revoice a line.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume