Meta AI disclosure: which Sume outputs count, video and audio
Meta asks for disclosure of photorealistic video and realistic audio that was digitally created or altered. A sorting guide for video, avatar and music.

Meta's manipulated media policy asks people to use its AI-disclosure tool when they post organic photorealistic video or realistic audio that was digitally created or altered. Penalties are possible, and Meta may add informative labels to high-risk deceptive content. Photorealistic generated video, talking avatars and realistic voice or music generally land in scope; clearly stylised animation is the grey area.
The policy points
This is the organic-content policy. Paid ads have a separate path in Ads Manager, covered in other posts.
| Item | Detail |
|---|---|
| Disclose | Organic photorealistic video or realistic audio, digitally created or altered |
| Mechanism | Meta's AI-disclosure tool when posting |
| Consequence | Penalties possible |
| Labels | Informative labels for high-risk deceptive content |
Sorting Sume outputs
The table is a working guide, not a legal ruling. It applies the policy's wording to the output types that Sume documents.
| Output | Docs | Lean |
|---|---|---|
| Photorealistic generated video | Video generation | Disclose |
| Talking avatar video of a realistic person | Avatar video | Disclose |
| Generated music with realistic instrumentation | Music Router | Disclose if posted as realistic audio |
| Clearly stylised or animated clip | Video generation | Judgement call; label when unsure |
| Captions burned onto your own clip | Video captions | No synthetic content added |
Make it a workflow step
Put the disclosure decision at the same moment as the upload. A per-clip sheet with the job ID, the output type and a yes or no for the toggle makes the choice visible. The video API returns a job ID on submit, and the avatar docs and Music Router docs cover the other two output types.
When a clip mixes real footage with a generated segment, consider disclosing the whole post. The policy summary we read says nothing about proportions, so label when unsure.
clip,job_id,output_type,meta_ai_toggle
spring-hero-9x16.mp4,<job id>,photorealistic video,yes
spring-theme.wav,<job id>,realistic audio,yes
spring-cartoon-cut.mp4,<job id>,stylised animation,reviewCaveats
Disclosure is cheap compared with a takedown or a trust problem. Leaning towards labelling is the lower-risk default.
- The policy page can change; check the date you read it.
- This post does not cover ads, which carry their own labels.
- If your output type is not in the table, ask which part a viewer would take as real.
Edge cases
Voice is the case people miss. A realistic synthetic voice over a real photo is still realistic audio that was digitally created, and falls under the same wording. A music bed generated for a clip belongs on the same list if it is presented as a real recording.
Another edge is partial edits. If you trim, caption or crop your own footage, you have not added synthetic content, and the table above says so. If you replace a face, a voice or a background with generated material, you have.
When two people disagree on a clip, ask which one would feel misled if they learned how it was made. Label if the answer is anyone.
Sources
Related posts
More in Use cases
- YouTube AI label: player or description, by how real it looks
YouTube shows the altered or synthetic label in the player for photorealistic content, in the expanded description for animated. What it means.
- YouTube end screens need a 25-second video: plan the clip length
YouTube end screens need videos 25 seconds or longer. Most Sume video models stop at 15 seconds, so here is how to reach 25 with long models or Timeline.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
Written by Sume