Full-duplex avatar listens while it speaks: script a clip instead
A full-duplex avatar hears you mid-sentence. If your content is scripted, build similar beats into one Sume avatar clip with silence scenes. Curl included.
A full-duplex avatar listens while it speaks, so it can nod, react or be interrupted without losing its place. A rendered clip cannot react to a live viewer, but if your message is scripted you can write those beats into the clip: a pause for a demo, a reaction shot, a line then a silence. In Sume Avatar 1.0 you do this with video_inputs that mix spoken scenes and voice.type: "silence" scenes in one request.
Tavus describes the live behaviour for Griffin-Lite, a research preview, this way: every sub-second mini-turn it decides whether to speak up, react or wait, and perception continues while the avatar speaks (Tavus: Griffin, read 2026-10-05).
What a script can borrow
Match each live behaviour to what a script can and cannot do. Anything that depends on what the viewer says next is out of reach for a rendered clip. Anything you can predict is easy to write.
The honest summary of the table is that a script borrows the timing of a live avatar and none of its awareness. A nod in a silent scene is a nod the writer decided on in advance, and it will happen whether or not a viewer would have wanted it. That is fine for a demo, an ad or a lesson with a fixed pace, and wrong for a service desk.
| Live behaviour (Tavus describes) | Rendered clip equivalent | Possible? |
|---|---|---|
| Nods while the user speaks | Silence scene with the avatar visible | Yes, as a fixed beat |
| Waits instead of speaking | voice.type silence, duration required | Yes |
| Speaks up at the right moment | Spoken scene ordered after a pause | Yes, at a fixed time |
| Is interrupted and recovers | Not applicable | No |
| Answers an unexpected question | Not applicable | No |
Three scenes in one clip
This request builds a hook, a four second silent beat and a call to action in one clip. A scene with type: silence needs a duration and must not contain a script; spoken scenes use type: text with script or input_text, not both. The total must fall between 4 and 60 seconds.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: beats-demo-001" \
-d '{
"avatar_handle": "presenter_01",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "Want to see it work?", "duration": 3},
"background": {"type": "prompt", "prompt": "Bright home office"}},
{"id": "beat", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Bright home office"}},
{"id": "cta", "voice": {"type": "text", "script": "Pick a template and drop in your photo.", "duration": 4},
"background": {"type": "prompt", "prompt": "Bright home office"}}
]
}'Limits to design around
Two limits from the docs shape the design. A final video has one avatar, and the scene backgrounds must resolve to one shared scene, so you cannot cut between locations in one request, as in the example where every scene uses the same background prompt. And estimated duration must be 4 to 60 seconds in total; a longer story needs several clips, joined afterwards. See Generate avatar video, read 2026-10-05.
A clip is also not live. It renders as a job, so you submit, poll or receive a webhook, and the viewer gets the finished video. That delay is fine for an ad or a tutorial and wrong for a call.
Beats cost the same as speech
Cost is per second of finished video, by quality tier, so a silent beat is billed like any other second of the clip. Plan beats as short as they can be. The price ladder lists the cost by length and tier. A four second silence at the standard tier is about as long as a nod, and costs the same per second as speech.
Which to pick
Use a full-duplex avatar when the viewer talks back, and a scripted clip when you know what you will say. For the longer comparison, see what full duplex means and when a clip is enough.
Sources
Related posts
More in Sume Avatar 1.0
- Full-duplex or script-driven avatar: six questions before you pick
Tavus Griffin-Lite is a full-duplex video conversation model in research preview. Six questions that separate it from Sume Avatar 1.0 talking video.
- Holiday ad captions on Avatar 1.0: slam, punch or tiktok-green?
Avatar 1.0 burns captions inline in slam (default), punch or tiktok-green. No separate billed caption job, and a caption failure keeps the clean video.
- Map avatar video scene ids to onboarding steps with scene_previews
Give each video_inputs scene a stable id and Sume returns it in scene_previews with start, end and duration, so one clip can drive per-step chapters.
- One voice across 23 languages: MAI-Voice-2.1 vs a Sume avatar voice
MAI-Voice-2.1 keeps one voice across 23 languages. A Sume voice has one primary language and a 409 guard on mismatch. What that means for avatars.
Written by Sume