A real person's photo avatar reading a script: YouTube disclosure
A Sume avatar made from a real person's photo and given a script can show them saying words they never said, which YouTube says needs a disclosure label.
If you build a Sume avatar from a real person's photo and give it a script, the video can show that person saying words they never said. YouTube's help page lists that case, making a real person appear to say or do something they did not, among the content creators must disclose.
What YouTube's page says
As read on 2026-10-08, YouTube requires disclosure when AI makes a real person appear to say or do something they did not do, or generates a realistic scene that did not occur. Labels appear in the video player for photorealistic content and in the expanded description for non-photorealistic content.
| Needs disclosure | Does not need disclosure |
|---|---|
| Making it appear someone gave advice they did not give | Cloning your own voice for voice-overs or dubs |
| Depicting a public figure doing something they did not do | Applying beauty filters |
| Realistic AI video of a match between real players | Non-realistic fantasy content |
How it maps to Sume avatars
Sume creates avatars from a prompt, structured props or a reference photo (photo, a public HTTPS image_url). A props or prompt avatar has no real person behind it. A photo avatar does, and the scripted speech is words that person may not have said.
The docs for creation and talking video describe inputs and limits; they do not set a consent or disclosure policy, so that decision stays with you and the platform where you publish.
- Photo of yourself, saying your own script: lowest risk, still consider a label.
- Photo of someone else: get their written permission before creating the avatar.
- Public figure: do not.
- Invented presenter from props: describe it as AI-generated in the post.
Practical steps
Keep a record of who agreed to what, keyed by the avatar handle. Put the disclosure where the platform asks for it, and add a short on-screen line if the clip could be mistaken for a real testimonial. Check the page above before publishing, since platform rules change.
Practical notes
Job lifecycle. Every Sume avatar request is job-backed. You submit, store the job id, and poll GET /v1/jobs/{id}/status until it completes, then read GET /v1/jobs/{id}/result. The shared lifecycle also lets you wait with sync or subscribe mode for a limited time instead of polling, and the result carries Sume-hosted artifacts rather than provider URLs, so you never handle a provider queue id. Plan your client around the job, not around a single blocking response.
Media inputs. Any media field, whether product_image, scene.image_url or an avatar image_url, must be a fetchable public HTTPS URL. Sume rejects localhost, private-network and non-HTTPS addresses before it submits anything, so a broken link fails fast instead of consuming a render. Host reference images where Sume can fetch them and keep them available until the job finishes.
Reading results. A completed result can include public media.sume.com video artifacts, and for avatar videos also public-safe preview fields such as preview_image_url and scene_previews. You can list or read finished videos with GET /v1/avatar-videos and GET /v1/avatar-videos/{id}. Store the job id and the video id together so a billing line can be traced back to a creative.
Sources
Related posts
More in Use cases
- Photographer portfolio reel from stills with Omni Flash
Turn 10 portfolio photos into a 60-second reel with Gemini Omni Flash: six seconds each, $7.50 at 720p on Sume, with 360p drafts first at $2.25.
- Pinterest video titles show 40 characters: burn the hook in
Pinterest shows the first 40 characters of a video ad title in the feed (30 for CJK) and hides the description. Burn the hook in with Sume captions.
- Pixel art and isometric video with Omni Flash: prompts
Prompt wording for pixel-art and isometric looks in Gemini Omni Flash, what Google's page says the model cannot control, and the cost of a draft on Sume.
- Podcast audio to a vertical video: cover stills plus a spine, $0.30
Turn 59 seconds of podcast audio into a captioned 1080x1920 video using cover-art stills as slots and the audio as the spine: render $0.10, captions $0.20.
Written by Sume