Hook, demo and CTA in one avatar video with a silence beat

Use video_inputs for a three-beat avatar video: a spoken hook, a silent product beat and a spoken CTA. One avatar, one shared scene, 4 to 60 seconds.

5 min readSume
All posts

Send ordered video_inputs to POST /v1/avatar-1.0/talking-video instead of one script. Each entry is a beat with a voice: type: text for speech, or type: silence with a required duration for a pause. The total planned duration must be 4 to 60 seconds, and all scenes share one avatar and one scene.

The three beats

A short ad usually has a hook, a moment where the product is on screen, and a call to action. In video_inputs that is two spoken entries and one silent entry. A silent beat has no speech, so script and input_text are not permitted on it.

Spoken entries use type: text with one of script or input_text, never both. Give each entry an id and a duration in seconds, and a background such as a prompt.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-video-multi-001" \
  -d '{
    "avatar_handle": "sume_clawra",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "video_inputs": [
      { "id": "hook", "voice": { "type": "text", "script": "Wait, this turned one selfie into a whole video?", "duration": 3 },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
      { "id": "demo", "voice": { "type": "silence", "duration": 4 },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
      { "id": "cta", "voice": { "type": "text", "input_text": "Pick a template, drop in your photo, and it builds the clip.", "duration": 5 },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }
    ]
  }'

The constraints

From docs.sume.com/models/avatar-videos, read 2026-10-05
ConstraintValue
Planned duration4 to 60 seconds inclusive
script and video_inputsSend one, not both
Avatars per final videoOne resolved avatar
Scene backgroundsMust resolve to one shared scene
aspect_ratio1:1, 3:4, 9:16 (default), 4:3, 16:9
resolution720p at this time
qualitystandard, plus (default) or max

What the silent beat is for

The silent beat is where the product or scene carries the message. Pair it with product_image or a photo scene (scene: { "type": "photo", "image_url": "https://..." }), and keep it short. Four seconds of nothing said feels long in a 12-second video.

Media fields must be fetchable public HTTPS URLs. Sume rejects localhost, private-network and non-HTTPS addresses before it submits.

Review before you render

Run the same body as a preview first. You get one still per scene in scene_previews[]. Later scene stills are pose-anchored continuations of the first frame when the scene is shared, so you see the continuity before you pay for the video.

Add captions with enabled: true and a style if you want them burned into the final video. Korean scripts need a Hangul style.

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume