Add a pause to an AI avatar video with a silence beat
To make a Sume avatar stop talking for a few seconds, use a silence scene in video_inputs: duration required, no script allowed. Example, limits and cost.
To add a pause to a Sume avatar video, replace the single script with an ordered video_inputs array and put a scene with voice.type: "silence" where the pause goes. A silence scene needs a duration in seconds and must not carry script or input_text. The avatar stays on screen without speaking, which suits a product reveal, a beat for a screen recording, or a gap before a call to action.
The silent seconds count toward the video. The planned total across all scenes must still land in the 4 to 60 second window, and you pay the per-second rate for the whole length.
The scene types
Each scene has a voice and a background.
| `voice.type` | Fields | Notes |
|---|---|---|
text | Exactly one of script or input_text; optional duration | A spoken scene |
silence | duration required; no script or input_text | A non-speaking beat |
A three-scene example
This plan speaks a hook, holds for four seconds, then closes. Every scene uses the same background prompt, because current execution resolves scene backgrounds to one shared scene and supports one avatar per final video.
{
"avatar_handle": "studio_presenter",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook",
"voice": {"type": "text", "script": "Here is the new dashboard.", "duration": 3},
"background": {"type": "prompt", "prompt": "Bright studio, clean desk"}},
{"id": "pause",
"voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Bright studio, clean desk"}},
{"id": "cta",
"voice": {"type": "text", "input_text": "Try it free today.", "duration": 3},
"background": {"type": "prompt", "prompt": "Bright studio, clean desk"}}
]
}What it costs
That plan is 10 seconds in total. At the plus rate of $0.245 per second with no product image it is $2.45, of which the four silent seconds are $0.98. Silence is not free, so cut it to what the edit needs. Confirm rates with GET /v1/catalog.
Alternatives
- If the pause is for B-roll, render the avatar speech as two clips and put the B-roll between them on a timeline.
- If you want captions over the speech, set
captionson the same request; captions follow the spoken text only. - To preview framing first, create an avatar video preview with the same
video_inputs; you get one still per scene.
Limits
A silence scene does not give you a different camera angle or a gesture; it is the same avatar and the same shared scene without speech. For a different scene, use a separate clip. Estimated duration over 60 seconds is rejected, so long pauses eat the budget quickly. A good rule is to keep silence under a third of the total length. If the point of the pause is to let a viewer read something on screen, check the on-screen time first: four seconds is long on a 9:16 short and short for a dense slide. Rendering a draft on standard ($0.184 per second) before a max final ($0.55 per second) costs about a third of the final and shows whether the rhythm works.
Sources
Related posts
More in Sume Avatar 1.0
- AI face swap video cost: the $2.76 to $8.25 ceiling per clip
Sume's Beta avatar face swap reserves a 15-second maximum: $2.76 on standard, $3.68 on plus, $8.25 on max. Constraints and when to use talking video instead.
- avatar-1.0/image-to-video is deprecated: switch to veed/fabric-1.0
Sume's avatar-1.0 image-to-video routes are deprecated aliases of VEED Fabric 1.0. Same body, new URL: audio_url, duration_seconds and one image source.
- Avatar 1.0 talking video: 4 to 60 seconds, cost by quality tier
Work out what a 4, 15, 30 or 60 second Avatar 1.0 talking video costs at standard, plus and max quality, using the per-second rates in the Sume catalog.
- Avatar handle with a leading @: how Sume stores and reuses it
Sume accepts an avatar_handle with or without a leading @ and stores it without. Why a stable handle beats a generated id, and how to reuse it.
Written by Sume