UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs
Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.
A UGC-style ad is three beats: a hook that someone says to the camera, a moment that shows the product, and a call to action. In the Sume Avatar 1.0 talking-video route, send those as three ordered video_inputs to POST /v1/avatar-1.0/talking-video: a voice of type: "text" for the hook, a voice of type: "silence" with a duration for the demo beat, and a spoken CTA. The total planned duration must fall in 4 to 60 seconds, and the route resolves one avatar and one shared scene for the whole video.
The request
The silence beat is the documented way to leave room for the product. It takes a required duration, and it does not accept script or input_text. Spoken scenes take exactly one of script or input_text, not both. Use the top-level avatar_handle to name a ready avatar, and send Idempotency-Key on every create.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-ad-serum-001" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "I stopped buying three serums after this one.", "duration": 4},
"background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}},
{"id": "demo", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}},
{"id": "cta", "voice": {"type": "text", "input_text": "Link in my bio, first bottle is half off today.", "duration": 5},
"background": {"type": "prompt", "prompt": "Bathroom mirror selfie framing, natural light"}}
]
}'What the route will and will not do
The constraints come from the docs. Read them before you design a storyboard around a change of place.
- One resolved avatar for each final video. A second speaker needs a second job.
- Scene backgrounds must resolve to one shared scene, so the three prompts above are the same on purpose. A new room for each beat is not supported by the current execution.
aspect_ratiosupports 1:1, 3:4, 9:16, 4:3 and 16:9, with a default of 9:16, andresolutionis 720p for now.- Media fields must be fetchable public HTTPS URLs. A
product_imageis optional, and a video without one is a productless avatar video. - Scripts that Sume estimates above 60 seconds must be shortened or split into several jobs.
Quality tier before you scale
quality takes standard, plus or max. plus is the default when you omit it, standard is the fastest path, and max is the highest tier with a slower turnaround. For an ad you will test, make the first version at standard, read the script on screen, and move to plus or max for the version you will run.
| Beat | voice.type | Required fields | Not allowed |
|---|---|---|---|
| Hook | text | script or input_text, duration | Both script and input_text |
| Demo | silence | duration | script, input_text |
| CTA | text | script or input_text, duration | Both script and input_text |
Put the product in the silent beat
A silent beat on an avatar does not show the product by itself. The docs give you two honest ways to fill it. First, set a product_image on the request so the avatar video includes the product. Second, generate the avatar video with the silence beat, then place product footage over that stretch in a Timeline compose, so the speaker's audio stays continuous. Both are in the docs, and the second one is the same pattern that the B-roll posts on this blog use. Whichever you pick, keep the silence beat short, because a long silent stretch on a talking head reads as a mistake.
Write the CTA in the same voice as the hook. A UGC ad that changes from first person to brand copy in the last scene loses the one thing it was making, which is the feel of a person talking. Run three variants of the hook with the same demo and CTA and compare them, using a separate idempotency key for each.
Sources
Related posts
More in Sume Avatar 1.0
- What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
- Can I use Tavus Griffin-Lite yet? Preview status and what to ship
Tavus Griffin-Lite is a research preview for select trusted testers, not open to customers. Here is what the page says, and what you can build on Sume today.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume