UGC hook, demo, CTA: one avatar video with a silence beat
Build a 9:16 UGC-style ad as three scenes in one avatar talking-video request: spoken hook, silent demo beat, spoken CTA, inside the 4-60 second window.
A UGC-style ad is three beats: a spoken hook, a demo where the person stops talking, and a spoken call to action. In Sume you send those beats as ordered video_inputs to POST /v1/avatar-1.0/talking-video, using a silence voice for the demo. The planned total must land between 4 and 60 seconds, and the default aspect ratio is 9:16.
Why scenes instead of one script
A single script makes the avatar talk for the whole clip. Real UGC does not work that way. The creator says the hook, shows the product while the room goes quiet, then talks again. video_inputs models that rhythm. Each scene has an id, a voice and a background.
Provide script or video_inputs, never both. The current execution supports one resolved avatar for each final video, and it expects the scene backgrounds to resolve to one shared scene, so keep the same bedroom or kitchen prompt on every scene.
The three scenes
The request below is the pattern from the avatar video docs. A spoken scene uses type: text with exactly one of script or input_text. A silent scene uses type: silence and needs only a duration.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-three-scene-001" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "Wait, this turned one selfie into a whole video?", "duration": 3},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
{"id": "demo", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
{"id": "cta", "voice": {"type": "text", "input_text": "Pick a template, drop in your photo, done.", "duration": 5},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}}
]
}'Quality and format knobs
quality accepts standard, plus and max. If you omit it, you get plus. standard is the fastest path, and max is the highest tier with a slower turnaround. Resolution is 720p at this time, and aspect_ratio accepts 1:1, 3:4, 9:16, 4:3 and 16:9.
| Setting | Allowed values | Default |
|---|---|---|
| quality | standard, plus, max | plus |
| aspect_ratio | 1:1, 3:4, 9:16, 4:3, 16:9 | 9:16 |
| resolution | 720p | 720p |
| planned duration | 4 to 60 seconds | none |
Writing a hook that fits three seconds
Duration is part of each voice entry, so write for the clock. A three-second hook is about eight spoken words. A five-second call to action fits a single sentence. Put the one claim the viewer must remember in the hook, and keep the demo beat silent so the product on screen does the talking.
If your total lands outside 4 to 60 seconds, Sume does not stretch it. Shorten the script or split the idea into two jobs and join them later with a Timeline render.
Run it and collect the file
Every avatar request creates a job, so you do not wait on the HTTP call. Send an Idempotency-Key that names the ad and the version, such as ugc-three-scene-001, and poll the job until it completes. If your network drops after the submit, send the same body with the same key and you get the same job back instead of a second charge.
When you test a new hook, change the key with the text. A key that never changes is a promise that the request never changes.
Add captions in the same call
UGC for TikTok, Reels and Shorts is watched with the sound off more often than not. The optional captions object burns styles into the clean final MP4 from the spoken text. Sume never burns captions into preview stills, so you can approve the look first. Read the avatar video docs for the caption fields.
Sources
Related posts
More in Sume Avatar 1.0
- UGC-style ad batch: ten 12-second hooks on Avatar 1.0 for $29.40
Ten 12-second UGC-style hook variants cost about $29.40 on Sume Avatar 1.0 plus, $22.08 on standard. The request, captions and checks.
- UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs
Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.
- Voiceover too long for H3 Max lip sync: the 5 to 14.8 s window
MiniMax H3 Max lip sync on Sume takes Sume-hosted audio of 5 to 14.8 seconds, refuses other lengths instead of clamping, and caps the file at 10 MB.
- What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
Written by Sume