Full-duplex or script-driven avatar: six questions before you pick
Tavus Griffin-Lite is a full-duplex video conversation model in research preview. Six questions that separate it from Sume Avatar 1.0 talking video.
Tavus Griffin-Lite, reported on 2026-10-01, is a full-duplex video conversation model, in research preview for invited testers (per a third-party tracker, read 2026-10-05; I read no more detail than that). Sume Avatar 1.0 talking video is script-driven: you send a script or an ordered list of scenes, and a job returns a finished clip. Before you pick, answer six questions. They tell you whether you need a live conversation, or a rendered clip.
The six questions
Question 1 is the one that decides the rest. If a person has to interrupt the avatar, a script cannot give you that. If the avatar says the same lines to everyone, a script gives you a file you can review, caption and reuse.
| Question | If the answer is yes | What Sume Avatar 1.0 does |
|---|---|---|
| 1. Does the viewer talk back while the avatar speaks? | You need a live, two-way model. | No. The input is a script or video_inputs. The result is a clip. |
| 2. Is the exact wording fixed before release? | A script is enough. | Yes. Provide script or video_inputs, not both. |
| 3. How long is the piece? | Over 60 seconds means several jobs. | Sume accepts a target duration of 4 to 60 seconds. |
| 4. How many people are on screen? | More than one means a different design. | One resolved avatar for each final video. |
| 5. What frame do you need? | Set it per channel. | aspect_ratio: 1:1, 3:4, 9:16 (default), 4:3, 16:9. resolution: 720p at this time. |
| 6. Can you review before paying for the render? | Previews help. | avatar-video-previews returns first-frame stills, then generate-video renders. |
What a script gives you that a call does not
A script-driven clip is a durable artifact. The result includes a public media.sume.com video, and may include preview_image_url and scene_previews. You can poll the job, read its events, and run the same text again with another quality tier: standard, plus (the default) or max.
You can also add silence beats. A scene with voice.type: "silence" has a required duration and no speech, so the avatar can hold a pause where a listener would nod. That is a scripted pause, not a reaction to a real person.
What a preview-only model leaves open
For Griffin-Lite, the only fact I can state is the research-preview status. Price, latency, resolution, API access and safety terms are things to ask the vendor. Do not fill them in from memory.
If a launch date matters to your plan, build the script-driven version first. The avatar, the script and the captions you make now will still be valid when a live option opens.
Decision in one line
- Need two-way conversation: wait for access, and ask the vendor the open questions.
- Need a finished clip for a feed, an ad or a lesson: use Avatar 1.0 now.
- Not sure: render one 15-second clip, read its preview stills first, and see whether a clip answers the need.
Sources
Related posts
More in Sume Avatar 1.0
- Holiday ad captions on Avatar 1.0: slam, punch or tiktok-green?
Avatar 1.0 burns captions inline in slam (default), punch or tiktok-green. No separate billed caption job, and a caption failure keeps the clean video.
- Map avatar video scene ids to onboarding steps with scene_previews
Give each video_inputs scene a stable id and Sume returns it in scene_previews with start, end and duration, so one clip can drive per-step chapters.
- One voice across 23 languages: MAI-Voice-2.1 vs a Sume avatar voice
MAI-Voice-2.1 keeps one voice across 23 languages. A Sume voice has one primary language and a 409 guard on mismatch. What that means for avatars.
- Pick a stock avatar by avoid_for and brand_safety_notes, not looks
Sume's avatar catalog returns profile metadata with best_for, avoid_for, brand_safety_notes and casting_notes. Read them before you cast a presenter for a clip.
Written by Sume