AI avatar mock interview: live interviewer or question clips
Tavus lists an Interviewer PAL as a live use case. Sume can render each interview question as a short avatar clip with a thinking pause. Where each one fits.
For an AI avatar mock interview you can use a live interviewer that reacts to your answers, or a set of recorded question clips you answer one at a time. Tavus's docs list an Interviewer PAL among four example use cases (read 2026-10-03), which is the live route. Sume covers the clip route: each question becomes a 4-60 second avatar video, optionally with a silent thinking beat, and your app collects the answers.
The difference is follow-up. A live interviewer can ask "why that choice?" about what you just said. A clip cannot.
What does a live AI interviewer give you?
It adapts. Tavus's pitch for Griffin names rehearsing difficult conversations as one of its target uses (press release via Yahoo Finance, read 2026-10-03), and the model is described as reacting while you are still speaking. That is gated, though: Griffin-Lite is available only to select trusted testers and not to customers (Tavus Griffin post). Today's customer-facing route is Tavus's conversation API, where a PAL joins a room with max_call_duration defaulting to 3,600 seconds (create-conversation reference).
Live rehearsal suits practising unscripted pressure. It suits less well when every candidate needs the same questions, delivered the same way, for fairness or review.
How do question clips work in Sume?
Render one clip per question. Each is a script through POST /v1/avatar-1.0/talking-video, which accepts 4-60 seconds, in 9:16, 16:9 or other supported ratios. To give the candidate a moment, use video_inputs: a spoken scene with the question, then a scene with voice.type: "silence" and a required duration. Current execution supports one avatar per final video and one shared scene background, so the interviewer looks the same throughout (Generate avatar video).
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: interview-q1-001" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "16:9",
"quality": "plus",
"video_inputs": [
{"id": "q", "voice": {"type": "text", "script": "Tell me about a project that did not go to plan, and what you changed.", "duration": 6},
"background": {"type": "prompt", "prompt": "Quiet office, neutral light"}},
{"id": "think", "voice": {"type": "silence", "duration": 8},
"background": {"type": "prompt", "prompt": "Quiet office, neutral light"}}
]
}'What does Sume not do for interviews?
It does not listen. Sume has no live session, so the avatar cannot react to an answer, detect that a candidate has stopped talking, or ask a follow-up. Your own app plays the clip, records the answer and decides what happens next.
| Need | Live interviewer | Question clips |
|---|---|---|
| Follow-up questions | Yes | No |
| Same questions for every candidate | Only if scripted into the PAL | By construction |
| Review wording before use | No | Yes, preview stills and script |
| Candidate pace | Real-time | Their own, with fixed thinking pauses |
| Availability | Tavus conversation API; Griffin gated | Any Sume API key |
How do I structure the question set?
Treat each question as its own job with its own id in your system, so one rewrite does not re-render the set. The 4-60 second window covers the whole clip, so a 6-second question plus an 8-second silence beat is a 14-second job; a long scenario question belongs in its own clip. Give the clips a shared avatar_handle, aspect ratio and scene prompt so the interviewer looks the same on every screen.
If candidates may watch on mute or in a noisy room, turn on inline captions, which burn the spoken question into the MP4 (Generate avatar video). Captions on the question also make the set easier to review by a hiring team, because the wording is on the frame.
Where does each approach go wrong?
A live interviewer can drift. If the PAL's knowledge or context differs between sessions, two candidates get different interviews, which matters if the practice is scored. Tavus lets you set conversational_context per conversation and fixed custom_greeting text (read 2026-10-03), but the follow-ups remain generated.
A clip set can feel stiff. Silence beats are fixed length, so a candidate who needs longer must pause the player, and nothing reacts to a half-finished answer. Neither failure is hidden by the technology, which is why the choice should follow the purpose: scored, repeatable screening points to clips; open practice with feedback points to a live agent.
How should I combine them?
Use clips for a structured first round, where fairness and repeatability matter, and a live agent for open practice. Run an avatar-video preview for each question before the full render so the framing is approved once (Avatar video previews).
- One clip per question, named by an id in your own system.
- A silence beat of fixed length as the thinking pause.
- Do not present the interviewer as a person.
- Keep the candidate's recording and consent handling in your app.
Sources
Related posts
More in Use cases
- AI avatar of a real person: release checklist before the photo
Before you turn an employee's or creator's photo into a Sume avatar, get a written release. A checklist for scope, term and revocation, matched to the API.
- AI avatar sales agent: live SDR or personalised clips?
Tavus builds live SDR avatars on a per-minute plan. Sume renders one 4-60 second avatar clip per lead from a script. How to pick, plus a Python loop.
- AI backing track generator: an instrumental to sing or play over
Generate an AI backing track by prompt: tempo, key, form, no vocals. Sume returns one mixed MP3, no stems or click track, so trim and check the take by ear.
- AI character series on Shorts: avoid the same situation each time
YouTube's inauthentic content policy flags characters in identical situations with the same outcomes. How to keep an AI character and vary the story in Sume.
Written by Sume