Conversation rehearsal with AI video: Griffin or recorded clips

Tavus lists rehearsing hard workplace talks for Griffin. A recorded version on Sume: one avatar per clip, with a silence beat where the learner replies.

5 min readSume
All posts

What Tavus lists

Tavus's Griffin page, read 2026-10-03, names several uses for a full-duplex video model: sales development, healthcare, interviews and recruiting, learning and development, tutoring, technical troubleshooting, and conversation rehearsal for difficult workplace discussions. The model watches expression, gaze and body language and decides when to speak, so a rehearsal partner can react to you.

Griffin-Lite is available only to select trusted testers, and Tavus says the wider model follows safety work. So a live rehearsal partner is not something to plan a launch on this week.

What a recorded rehearsal can do

A practice session does not always need a partner that listens. Many courses use a fixed prompt: the other person says a line, the learner answers aloud or in writing, then sees a model answer. That shape works as video on Sume.

Sume's avatar video takes video_inputs, an ordered list of scenes. A scene can speak text or hold a silence beat, and a silence beat has a required duration and no text. Use a spoken scene for the colleague's line, then a silence beat for the learner's reply, then the next line.

{
  "avatar_handle": "team_lead",
  "aspect_ratio": "16:9",
  "video_inputs": [
    {"id": "line", "voice": {"type": "text", "script": "I need to talk about the missed deadline.", "duration": 4},
     "background": {"type": "prompt", "prompt": "Quiet office, soft daylight"}},
    {"id": "your-turn", "voice": {"type": "silence", "duration": 8},
     "background": {"type": "prompt", "prompt": "Quiet office, soft daylight"}}
  ]
}

Limits to plan around

  • One avatar per final video. Sume's docs say current execution resolves one avatar per video and expects the scene backgrounds to share one scene. A two-person role-play is two clips or one clip with one visible speaker.
  • 4 to 60 seconds per job. A long scenario is a chain of short clips, and the learner presses play for each.
  • The silence is fixed. The learner cannot interrupt, and the clip does not react to their answer.
  • Preview the first frame before spending on a full render, using avatar video previews.

A fair split

Use recorded clips for scripted practice you can review and caption, and keep a note to revisit live rehearsal when a full-duplex model is open to customers. Tell learners which side is a recording.

A course outline that fits

A short module could run as three clips: the setup, the difficult line with an eight-second pause, and a model response. Each clip is its own job, which keeps every file well inside the 60-second cap and lets you swap one line without re-rendering the rest.

Add captions with the captions option so the exercise works with sound off, and tell learners clearly that the other person is an AI avatar.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume