Full-duplex avatar listens while it speaks: script a clip instead

A full-duplex avatar hears you mid-sentence. If your content is scripted, build similar beats into one Sume avatar clip with silence scenes. Curl included.

4 min readSume
All posts

A full-duplex avatar listens while it speaks, so it can nod, react or be interrupted without losing its place. A rendered clip cannot react to a live viewer, but if your message is scripted you can write those beats into the clip: a pause for a demo, a reaction shot, a line then a silence. In Sume Avatar 1.0 you do this with video_inputs that mix spoken scenes and voice.type: "silence" scenes in one request.

Tavus describes the live behaviour for Griffin-Lite, a research preview, this way: every sub-second mini-turn it decides whether to speak up, react or wait, and perception continues while the avatar speaks (Tavus: Griffin, read 2026-10-05).

What a script can borrow

Match each live behaviour to what a script can and cannot do. Anything that depends on what the viewer says next is out of reach for a rendered clip. Anything you can predict is easy to write.

The honest summary of the table is that a script borrows the timing of a live avatar and none of its awareness. A nod in a silent scene is a nod the writer decided on in advance, and it will happen whether or not a viewer would have wanted it. That is fine for a demo, an ad or a lesson with a fixed pace, and wrong for a service desk.

Live avatar behaviours and a rendered equivalent, read 2026-10-05
Live behaviour (Tavus describes)Rendered clip equivalentPossible?
Nods while the user speaksSilence scene with the avatar visibleYes, as a fixed beat
Waits instead of speakingvoice.type silence, duration requiredYes
Speaks up at the right momentSpoken scene ordered after a pauseYes, at a fixed time
Is interrupted and recoversNot applicableNo
Answers an unexpected questionNot applicableNo

Three scenes in one clip

This request builds a hook, a four second silent beat and a call to action in one clip. A scene with type: silence needs a duration and must not contain a script; spoken scenes use type: text with script or input_text, not both. The total must fall between 4 and 60 seconds.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: beats-demo-001" \
  -d '{
    "avatar_handle": "presenter_01",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "video_inputs": [
      {"id": "hook", "voice": {"type": "text", "script": "Want to see it work?", "duration": 3},
       "background": {"type": "prompt", "prompt": "Bright home office"}},
      {"id": "beat", "voice": {"type": "silence", "duration": 4},
       "background": {"type": "prompt", "prompt": "Bright home office"}},
      {"id": "cta", "voice": {"type": "text", "script": "Pick a template and drop in your photo.", "duration": 4},
       "background": {"type": "prompt", "prompt": "Bright home office"}}
    ]
  }'

Limits to design around

Two limits from the docs shape the design. A final video has one avatar, and the scene backgrounds must resolve to one shared scene, so you cannot cut between locations in one request, as in the example where every scene uses the same background prompt. And estimated duration must be 4 to 60 seconds in total; a longer story needs several clips, joined afterwards. See Generate avatar video, read 2026-10-05.

A clip is also not live. It renders as a job, so you submit, poll or receive a webhook, and the viewer gets the finished video. That delay is fine for an ad or a tutorial and wrong for a call.

Beats cost the same as speech

Cost is per second of finished video, by quality tier, so a silent beat is billed like any other second of the clip. Plan beats as short as they can be. The price ladder lists the cost by length and tier. A four second silence at the standard tier is about as long as a nod, and costs the same per second as speech.

Which to pick

Use a full-duplex avatar when the viewer talks back, and a scripted clip when you know what you will say. For the longer comparison, see what full duplex means and when a clip is enough.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume