Photo avatar of a real person: when YouTube's AI label applies

A photo avatar scripted to say new words falls under YouTube's 'real person says something they did not' test. How Sume's three avatar inputs map to it.

3 min readSume
All posts

An avatar made from a photo of a real person, then scripted to say words that person never said, matches the first disclosure trigger on YouTube's help page: making a real person appear to say or do something they did not do. If you upload that video to YouTube, use the altered or synthetic content disclosure. An avatar made from a text prompt or profile traits depicts nobody in particular, and the page does not settle whether a photoreal one needs the label.

The three Sume avatar inputs

Sume Avatar 1.0 creates an avatar from a prompt, from structured profile traits (the props type) or from a reference image (the photo type). Creation costs $0.95 per avatar (catalog, read 2026-10-09). The photo input needs a public HTTPS image URL; the docs check the URL for safety and for being an image, and say nothing about whose face it is.

The same reading applies to a before-and-after: if a person appears on camera saying their own words and you only change the language, the page lists dubs in a cloned own voice as exempt, but it describes voice rather than a re-rendered face. Where the face is generated and the words are new, treat it as altered content.

Avatar inputs against YouTube's trigger (YouTube Help and Sume docs, read 2026-10-09)
Sume inputDepicts a real person?YouTube test
PromptOnly if you describe oneJudgment call on realism
Profile traits (props)No specific personJudgment call on realism
Photo (image_url)Yes, if the photo is of oneReal person saying new words: disclose

What to record

Sume's docs do not describe a consent field on avatar creation, so the record lives in your system. For each avatar_handle, keep who the person is, what they agreed to, the date, and the job ids of every video made with that handle. Handles are normalized without the leading @, which makes them a clean key.

YouTube's page exempts non-realistic content and minor edits such as beauty filters. A talking, photoreal avatar of a real person is neither.

Because the label lives on YouTube's side, nothing in the Sume job changes if you add it. Do it at upload: the disclosure setting is part of the upload flow in YouTube Studio. Keep a copy of the label decision with the job id so the choice is reviewable.

  • Photo of yourself, scripted by you, is still a re-voiced face: decide and label accordingly.
  • Dubbing in your own cloned voice is the page's example of no disclosure, but that is about voice only.
  • When unsure, label it. The label costs a checkbox.

A cheap pre-publish check

Before upload, preview the first frames through the avatar video previews route, confirm the person pictured is the person who agreed, and only then render. Previews are stills, so a mistake costs far less than a full render at $0.245 per second at plus.

If you work with talent, put the permission in writing before the photo goes into the request. A short agreement should say what the person allows (a talking avatar made from this photo), where it may appear (platforms and paid placements), for how long, and how they can withdraw. Sume's endpoint cannot verify that agreement, so the safest place for the record is next to the avatar_handle in your own system. The $0.95 creation fee is negligible next to the cost of a takedown.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume