Put the AI disclosure in scene one: Sume avatar video_inputs recipe
A copyable Sume Avatar 1.0 request with a 3-second disclosure scene and a 12-second message. The 15 seconds cost $2.76 at standard, $3.68 at plus.
Short answer
Use ordered video_inputs on the talking-video request and make the first scene the disclosure line. The disclosure is then spoken by the presenter at the start of every clip, it is part of the video file, and inline captions can show it as text. The request below has a 3-second disclosure scene and a 12-second message, 15 seconds in total, which costs $2.76 at standard, $3.68 at plus and $8.25 at max.
The request
This is a multi-scene request. Each scene has an id, a voice of type text with a script and a duration, and a background. The docs say the current execution expects the backgrounds to resolve to one shared scene, so keep the same background prompt in both scenes.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: disclosed-clip-001" \
-d '{
"avatar_handle": "product_host",
"aspect_ratio": "9:16",
"quality": "plus",
"captions": { "enabled": true, "style": "slam" },
"video_inputs": [
{
"id": "disclosure",
"voice": { "type": "text", "script": "I am an AI presenter.", "duration": 3 },
"background": { "type": "prompt", "prompt": "Bright studio desk" }
},
{
"id": "message",
"voice": { "type": "text", "script": "Version two ships Friday with faster exports.", "duration": 12 },
"background": { "type": "prompt", "prompt": "Bright studio desk" }
}
]
}'What it costs
Cost follows total planned seconds times the tier rate. The disclosure scene is 3 of the 15 seconds, so it is a fifth of the bill.
| Tier | Rate per second | 15 seconds | Disclosure scene (3 s) |
|---|---|---|---|
| standard | $0.184 | $2.76 | $0.55 |
| plus | $0.245 | $3.68 | $0.74 |
| max | $0.55 | $8.25 | $1.65 |
Notes on the fields
The total planned duration must stay inside 4 to 60 seconds. Spoken scenes use type text with one of script or input_text, not both. A beat with no speech uses type silence and requires a duration. Captions are optional, are applied to the final MP4 only, and a failure in that stage can leave the video fine with captions.status set to failed, so check the result. If the language is Korean, pick a Hangul caption style.
- Use the same fixed sentence in every clip.
- Keep the disclosure short enough to fit its duration.
- Run it through a preview first if a reviewer must approve.
What this does and does not do
It makes the presenter say that it is an AI presenter. It does not set the YouTube label, which you do at upload. The YouTube help page says disclosing does not limit reach or monetization eligibility, so repeating the line costs seconds of video rather than audience.
If you make many clips, keep the request body in a template and fill in only the message script and duration. Count words before you send: a 12-second scene fits a short sentence or two, and a script that Sume estimates as longer than its planned time is a reason to cut words rather than raise the duration. Long messages are better split across jobs, each with its own disclosure scene, than stretched past the 60-second ceiling.
Variations
A silence beat after the disclosure can give viewers a moment before the message starts; it costs the same per second as speech, so a 1-second beat at plus is $0.245. For a vertical short use 9:16, which is the default; for slides use 16:9. If a reviewer wants to see the look before you pay for the video, create it as a preview first and render from the preview id.
Sources
Related posts
More in Sume Avatar 1.0
- Video call avatar or scripted talking video: which do you need?
A real-time avatar answers people live; a scripted talking video is a file you render and review first. How to choose, and what Sume Avatar 1.0 covers.
- Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?
Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.
- Reuse one AI spokesperson across videos with an avatar handle
A Sume avatar_handle is a stable name for a ready avatar. Sume strips a leading @, so you create once and call the handle in every talking-video request.
- Add a silent beat to an AI avatar video with voice type silence
In Sume multi-scene avatar videos a scene with voice type silence is a pause with no speech. It needs a duration and rejects script text. Rules and example.
Written by Sume