Hook, demo and CTA in one avatar video with a silence beat
Use video_inputs for a three-beat avatar video: a spoken hook, a silent product beat and a spoken CTA. One avatar, one shared scene, 4 to 60 seconds.
Send ordered video_inputs to POST /v1/avatar-1.0/talking-video instead of one script. Each entry is a beat with a voice: type: text for speech, or type: silence with a required duration for a pause. The total planned duration must be 4 to 60 seconds, and all scenes share one avatar and one scene.
The three beats
A short ad usually has a hook, a moment where the product is on screen, and a call to action. In video_inputs that is two spoken entries and one silent entry. A silent beat has no speech, so script and input_text are not permitted on it.
Spoken entries use type: text with one of script or input_text, never both. Give each entry an id and a duration in seconds, and a background such as a prompt.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-video-multi-001" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{ "id": "hook", "voice": { "type": "text", "script": "Wait, this turned one selfie into a whole video?", "duration": 3 },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
{ "id": "demo", "voice": { "type": "silence", "duration": 4 },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
{ "id": "cta", "voice": { "type": "text", "input_text": "Pick a template, drop in your photo, and it builds the clip.", "duration": 5 },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }
]
}'The constraints
| Constraint | Value |
|---|---|
| Planned duration | 4 to 60 seconds inclusive |
script and video_inputs | Send one, not both |
| Avatars per final video | One resolved avatar |
| Scene backgrounds | Must resolve to one shared scene |
aspect_ratio | 1:1, 3:4, 9:16 (default), 4:3, 16:9 |
resolution | 720p at this time |
quality | standard, plus (default) or max |
What the silent beat is for
The silent beat is where the product or scene carries the message. Pair it with product_image or a photo scene (scene: { "type": "photo", "image_url": "https://..." }), and keep it short. Four seconds of nothing said feels long in a 12-second video.
Media fields must be fetchable public HTTPS URLs. Sume rejects localhost, private-network and non-HTTPS addresses before it submits.
Review before you render
Run the same body as a preview first. You get one still per scene in scene_previews[]. Later scene stills are pose-anchored continuations of the first frame when the scene is shared, so you see the continuity before you pay for the video.
Add captions with enabled: true and a style if you want them burned into the final video. Korean scripts need a Hangul style.
Related posts
More in Sume Avatar 1.0
- Avatar inline captions vs standalone video captions: what is billed
Inline captions on a Sume avatar video create no separate caption job. Standalone captions are a billed job on a public URL. When each one is right.
- Avatar video aspect ratios: five choices, 720p only, 9:16 default
Sume Avatar 1.0 accepts 1:1, 3:4, 9:16, 4:3 and 16:9, defaults to 9:16, and renders 720p only. What each choice means for a vertical or landscape ad.
- Avatar package with captions and soundtrack: which video_url you get
In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.
- Avatar video with a product image: the premium is 1-3 cents a second
A product_image on a Sume Avatar 1.0 video adds $0.010 (standard), $0.013 (plus) or $0.030 (max) per second: 30 to 90 cents on a 30-second ad.
Written by Sume