Put the AI disclosure in scene one: Sume avatar video_inputs recipe

A copyable Sume Avatar 1.0 request with a 3-second disclosure scene and a 12-second message. The 15 seconds cost $2.76 at standard, $3.68 at plus.

5 min readSume
All posts

Short answer

Use ordered video_inputs on the talking-video request and make the first scene the disclosure line. The disclosure is then spoken by the presenter at the start of every clip, it is part of the video file, and inline captions can show it as text. The request below has a 3-second disclosure scene and a 12-second message, 15 seconds in total, which costs $2.76 at standard, $3.68 at plus and $8.25 at max.

The request

This is a multi-scene request. Each scene has an id, a voice of type text with a script and a duration, and a background. The docs say the current execution expects the backgrounds to resolve to one shared scene, so keep the same background prompt in both scenes.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: disclosed-clip-001" \
  -d '{
    "avatar_handle": "product_host",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "captions": { "enabled": true, "style": "slam" },
    "video_inputs": [
      {
        "id": "disclosure",
        "voice": { "type": "text", "script": "I am an AI presenter.", "duration": 3 },
        "background": { "type": "prompt", "prompt": "Bright studio desk" }
      },
      {
        "id": "message",
        "voice": { "type": "text", "script": "Version two ships Friday with faster exports.", "duration": 12 },
        "background": { "type": "prompt", "prompt": "Bright studio desk" }
      }
    ]
  }'

What it costs

Cost follows total planned seconds times the tier rate. The disclosure scene is 3 of the 15 seconds, so it is a fifth of the bill.

Cost of a 3 + 12 second clip (rates as of 2026-10-08)
TierRate per second15 secondsDisclosure scene (3 s)
standard$0.184$2.76$0.55
plus$0.245$3.68$0.74
max$0.55$8.25$1.65

Notes on the fields

The total planned duration must stay inside 4 to 60 seconds. Spoken scenes use type text with one of script or input_text, not both. A beat with no speech uses type silence and requires a duration. Captions are optional, are applied to the final MP4 only, and a failure in that stage can leave the video fine with captions.status set to failed, so check the result. If the language is Korean, pick a Hangul caption style.

  • Use the same fixed sentence in every clip.
  • Keep the disclosure short enough to fit its duration.
  • Run it through a preview first if a reviewer must approve.

What this does and does not do

It makes the presenter say that it is an AI presenter. It does not set the YouTube label, which you do at upload. The YouTube help page says disclosing does not limit reach or monetization eligibility, so repeating the line costs seconds of video rather than audience.

If you make many clips, keep the request body in a template and fill in only the message script and duration. Count words before you send: a 12-second scene fits a short sentence or two, and a script that Sume estimates as longer than its planned time is a reason to cut words rather than raise the duration. Long messages are better split across jobs, each with its own disclosure scene, than stretched past the 60-second ceiling.

Variations

A silence beat after the disclosure can give viewers a moment before the message starts; it costs the same per second as speech, so a 1-second beat at plus is $0.245. For a vertical short use 9:16, which is the default; for slides use 16:9. If a reviewer wants to see the look before you pay for the video, create it as a preview first and render from the preview id.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume