UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs

Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.

5 min readSume
All posts

A UGC-style ad is three beats: a hook that someone says to the camera, a moment that shows the product, and a call to action. In the Sume Avatar 1.0 talking-video route, send those as three ordered video_inputs to POST /v1/avatar-1.0/talking-video: a voice of type: "text" for the hook, a voice of type: "silence" with a duration for the demo beat, and a spoken CTA. The total planned duration must fall in 4 to 60 seconds, and the route resolves one avatar and one shared scene for the whole video.

The request

The silence beat is the documented way to leave room for the product. It takes a required duration, and it does not accept script or input_text. Spoken scenes take exactly one of script or input_text, not both. Use the top-level avatar_handle to name a ready avatar, and send Idempotency-Key on every create.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ugc-ad-serum-001" \
  -d '{
    "avatar_handle": "sume_clawra",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "video_inputs": [
      {"id": "hook", "voice": {"type": "text", "script": "I stopped buying three serums after this one.", "duration": 4},
       "background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}},
      {"id": "demo", "voice": {"type": "silence", "duration": 4},
       "background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}},
      {"id": "cta", "voice": {"type": "text", "input_text": "Link in my bio, first bottle is half off today.", "duration": 5},
       "background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}}
    ]
  }'

What the route will and will not do

The constraints come from the docs. Read them before you design a storyboard around a change of place.

  • One resolved avatar for each final video. A second speaker needs a second job.
  • Scene backgrounds must resolve to one shared scene, so the three prompts above are the same on purpose. A new room for each beat is not supported by the current execution.
  • aspect_ratio supports 1:1, 3:4, 9:16, 4:3 and 16:9, with a default of 9:16, and resolution is 720p for now.
  • Media fields must be fetchable public HTTPS URLs. A product_image is optional, and a video without one is a productless avatar video.
  • Scripts that Sume estimates above 60 seconds must be shortened or split into several jobs.

Quality tier before you scale

quality takes standard, plus or max. plus is the default when you omit it, standard is the fastest path, and max is the highest tier with a slower turnaround. For an ad you will test, make the first version at standard, read the script on screen, and move to plus or max for the version you will run.

Avatar talking-video fields for a three-beat ad (Sume docs, read 2026-10-05)
Beatvoice.typeRequired fieldsNot allowed
Hooktextscript or input_text, durationBoth script and input_text
Demosilencedurationscript, input_text
CTAtextscript or input_text, durationBoth script and input_text

Put the product in the silent beat

A silent beat on an avatar does not show the product by itself. The docs give you two honest ways to fill it. First, set a product_image on the request so the avatar video includes the product. Second, generate the avatar video with the silence beat, then place product footage over that stretch in a Timeline compose, so the speaker's audio stays continuous. Both are in the docs, and the second one is the same pattern that the B-roll posts on this blog use. Whichever you pick, keep the silence beat short, because a long silent stretch on a talking head reads as a mistake.

Write the CTA in the same voice as the hook. A UGC ad that changes from first person to brand copy in the last scene loses the one thing it was making, which is the feel of a person talking. Run three variants of the hook with the same demo and CTA and compare them, using a separate idempotency key for each.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume