Add a silent beat to an AI avatar video with voice type silence
In Sume multi-scene avatar videos a scene with voice type silence is a pause with no speech. It needs a duration and rejects script text. Rules and example.
To put a pause in a Sume avatar video, add a scene to video_inputs with voice type silence and a duration in seconds. The scene has no speech. Duration is required, and script or input_text are not permitted on a silence scene. Spoken scenes keep using type text with exactly one of script or input_text.
The scene fields
From Generate avatar video, read 2026-10-08:
| voice.type | Required | Not allowed |
|---|---|---|
| text | duration plus one of script or input_text | Both script and input_text together |
| silence | duration | script, input_text |
Where a silent beat helps
A beat of quiet is useful when something on screen should carry the moment: a product being shown, a number appearing, or a caption landing. It also gives a hook room. A UGC-style clip might open with a spoken question, hold for a few seconds on the demo, then close with a call to action.
- Hook: spoken, 2-4 seconds.
- Demo: silence, enough time to show the product.
- Call to action: spoken, the final line.
Constraints that still apply
Silence counts toward the total. The planned duration of all scenes together must stay in the 4-60 second window. The execution supports one resolved avatar per final video, and it expects every scene background to resolve to one shared scene, so a silence beat does not give you a new location.
Preview the result before a full render. Preview creates one first-frame still per scene, and later stills are pose-anchored continuations of the first frame. Because a silence scene still gets a still, you can check that the avatar holds a natural pose through the pause. Details are on Avatar video previews.
Example body
Three scenes with a silent middle, shortened from the docs example.
{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"video_inputs": [
{"id": "hook",
"voice": {"type": "text", "script": "Wait, one photo became a video?", "duration": 3},
"background": {"type": "prompt", "prompt": "Casual bedroom framing"}},
{"id": "demo",
"voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Casual bedroom framing"}}
]
}Timing a pause
A silence scene is plain time. Four seconds feels long on camera, so use it only when something visibly happens: a product turning, a result appearing, or a step on a screen. For a pure beat of emphasis, one to two seconds is usually enough, but remember the clip's total must be at least four seconds.
Because every scene shares one background, a pause cannot move the speaker to another location. If you need a different setting, make a separate video and join the clips afterward. This keeps the avatar's identity and the scene consistent inside each render.
Captions follow the spoken text only, so silence scenes carry no caption words. Viewers will see a gap, which is usually what you want, but check the preview and the final result on a phone before you publish.
Sources
Related posts
More in Sume Avatar 1.0
- Sume Avatar 1.0 ratios vs TikTok ad ratios: three overlap, two do not
Avatar 1.0 makes 1:1, 3:4, 9:16, 4:3 or 16:9 at 720p, 4 to 60 s. TikTok's non-Spark ad page lists 9:16, 16:9 and 1:1, so 3:4 and 4:3 need a crop.
- Does Sume Avatar 1.0 speak Spanish or Korean? English only in code
The Avatar 1.0 talking-video prompt on main says English only. What that means for Spanish or Korean lines, and the audio-driven route to test instead.
- Synthesia needs a live consent clip: what do Sume photo avatars need?
Synthesia says personal avatars need a live consent recording. Sume's photo avatar takes an HTTPS image URL, so your own release process must cover consent.
- Synthesia Pro yearly, $64/mo: break-even vs Sume is 69.6 minutes
Synthesia lists Pro yearly at $64 a month for about 720 minutes a year. Sume standard matches that $768 outlay at about 69.6 minutes of video; plus at 52.2.
Written by Sume