AI listing video disclosure line: put it in the Sume output schema
Carry the AI disclosure and the original-photos link with each listing video: a typed output_schema field, what input does not carry, and a check.

To keep the AI disclosure attached to a listing video, ask the Format run for it as a field of the typed result. Bind an output_schema with a video file and a disclosure string, and store the whole object next to the listing. The video then never reaches your database without its disclosure line.
This matches the test in HousingWire's June 2026 article (read 2026-10-06): its fourth question asks whether the buyer will actually see the disclosure, and its fifth asks whether the agent can show what was real and what was changed. A required schema field helps with the fourth, not with the fifth. It cites California AB 723, effective 2026-01-01.
What does the run return, and what does it not?
The structured-output page says input and output_schema are different things. input is caller data that goes in; output_schema is the contract for what comes out. Sume builds output from what the run made and said, so a value you sent, such as a listing id or an MLS number, does not come back unless the run repeats it. Keep your identifiers on your side, keyed by the run id or your Idempotency-Key.
So the disclosure text has to be produced by the run. Give the wording in instruction, and ask for it in the schema.
| Field | Type | Who fills it |
|---|---|---|
video | SumeMediaFile# | The run, from the generated clip |
disclosure | string | The run, repeating your wording |
listing_id | string | Not returned; keep it in your own table by run id |
Request body
The shape is the one in Create a run: instruction, input, output_schema, primary_output_key, a cap, and Idempotency-Key. Replace acme/listing-tour with a Format you own; the catalog has no listing Format.
curl -sS -X POST "https://api.sume.com/v1/formats/acme/listing-tour/runs" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: listing-1042-v1" \
-d '{
"instruction": "Make a 20 second tour from the attached photos. Set disclosure to exactly: Video was created from listing photos using AI. Camera movement is simulated.",
"input": { "photos": ["https://cdn.example.com/1042/kitchen.jpg"] },
"output_schema": {
"name": "acme/listing-tour/v1", "strict": true,
"schema": { "type": "object", "additionalProperties": false,
"required": ["video", "disclosure"],
"properties": { "video": { "$ref": "SumeMediaFile#" }, "disclosure": { "type": "string" } } }
},
"primary_output_key": "video",
"generation_spend_cap_usd": 10
}'How do you check it?
A completed run can still carry output_error when the projection did not match your schema, so read output_error before output. Then compare output.disclosure with your wording in code and refuse to publish on a mismatch. The model wrote that string; your code should not assume it copied it exactly.
Use the listing id and a version as the Idempotency-Key, as in listing-1042-v1. The docs say the same key with the same body returns the original run with idempotency_hit: true and no second charge, and that you should bump the version only when you want a re-run. A time-based key would make every retry a new paid run.
The schema does not make the video lawful or accurate. It makes the disclosure a required field, which is the part software can enforce.
Sources
Related posts
More in Use cases
- AI video generator for real estate: which Sume Format?
Sume's 27 catalog Formats include none for property listings. What to call instead for a listing video, and the disclosure test agents face.
- Real estate listing video with an AI avatar: fair housing script check
Describe the home, not the buyer. A fair housing word check in Python, a listing-photo scene on Sume Avatar 1.0, and the cost of a 45-second listing clip.
- Reddit 15-second ad: join hook, demo and offer, then plan first
Build a 15-second ad from three clips on Timeline 1.0 and check its length and cost with the unbilled /plan call before you pay for a render.
- Reddit 15-second engaged views: does your voiceover CTA land?
Reddit's Engaged Video Views beta bills videos over 15 s at 15 s. Use TTS word timestamps to check that your call to action is spoken before second 15.
Written by Sume