China's implicit AI label: provider name and content ID in metadata

China's implicit AI label is metadata with the provider name and a content ID. Why a re-encode can drop it, and how to check a delivered file with ffprobe.

5 min readSume
All posts

China's implicit AI label is metadata embedded in the generated file that carries the service provider's name and a content ID, alongside the visible explicit label. It does not follow a file through every edit by itself: if a tool re-encodes the video, whether the metadata survives depends on that tool, so check the file you deliver.

The description of the label comes from Inside Privacy's summary of the measures (read 2026-10-02), which gives 1 September 2025 as the effective date. I did not read the national technical standard itself, so I cannot tell you which container field the standard names.

What is an implicit label?

Inside Privacy contrasts the two label types. An explicit label is a visible indicator in text, audio or graphics, placed at an appropriate location in the content. An implicit label is metadata embedded in the file, with the provider's name and a content ID.

For distribution platforms the summary describes a detection step. If a platform detects an implicit label, it adds a clear AI-generated indicator; user reports and other evidence lead to softer 'possibly' or 'suspected' labels. The metadata is therefore what lets a platform label a file without guessing.

Explicit versus implicit label, per Inside Privacy (read 2026-10-02)
LabelFormWho reads it
ExplicitVisible text, audio or graphicThe viewer
ImplicitMetadata: provider name, content IDPlatform detection

Why would a trim or a render lose it?

Container metadata lives next to the video stream, not inside the pixels. A tool that decodes and re-encodes the stream writes a new file, and whether it copies the old tags is a choice that tool makes.

Sume's server compiles ffmpeg on its worker for Video trim and returns a new MP4 for the range you ask for, and Timeline 1.0 assembles an audio spine plus ordered clips into one MP4. The docs describe the output as a new file and make no promise about metadata carried over from a source. Sume's docs describe what each tool returns and say nothing about a watermark or embedded provenance record on outputs, so treat the output as unmarked until you have checked the delivered file yourself.

How do I check a delivered file?

Run ffprobe on the exact file you will upload and read the format and stream tags. The script below prints them and exits with an error if ffprobe is not installed or the file is unreadable. An empty tag list means no provenance metadata is present in those fields.

Sume's Video inspect route also returns a probe for a hosted clip, but its documented fields are probe facts, stills and an optional transcript, so use ffprobe when you need the raw tags.

import json, subprocess, sys

def tags(path):
    out = subprocess.run(
        ["ffprobe", "-v", "error", "-print_format", "json",
         "-show_format", "-show_streams", path],
        capture_output=True, text=True, check=True,
    ).stdout
    info = json.loads(out)
    found = {"format": info["format"].get("tags", {})}
    for i, s in enumerate(info["streams"]):
        found["stream%d" % i] = s.get("tags", {})
    return found

if __name__ == "__main__":
    print(json.dumps(tags(sys.argv[1]), indent=2))

What does the visible label add?

The explicit label is the part a re-encode cannot remove, because it is drawn into the frames. That is why a two-layer approach is the sturdy one: a burned-in notice for the viewer, plus metadata for platform detection that you re-apply after the last export.

The two layers fail differently. Lose the metadata and the platform may fall back to user reports or its own detection. Lose the visible notice and the viewer sees nothing, whatever the platform decides. Neither layer substitutes for the other in the summary I read, which lists both as obligations rather than alternatives.

One more practical point: a file passes through several hands. A clip generated in one tool, trimmed in another and uploaded from a phone has had at least three chances to lose tags. Test the file in the form the platform actually receives, not the master you exported at the start.

What should you do about it?

Sume does not write a provider name or content ID into a file and does not offer a metadata-writing step in the docs I read. If you need one, add it in your own post-processing and verify it with the script above.

  • Decide who owns the implicit label: the generating tool, your own pipeline, or both.
  • Write the metadata last, after every trim, resize and composite, so no later step overwrites it.
  • Check the file on the platform side: re-upload a test clip and see what the platform shows.
  • Keep the visible explicit label in the pixels, because it survives a re-encode where tags may not.
  • Log the job id and the delivered artifact next to the content ID you assign.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume