Pre-render top 20 support answers as avatar clips; agent for the rest

Render the 20 most common answers as Sume avatar clips, serve them by intent, and send the rest to a person. Tavus Griffin is not open to customers.

4 min readSume
All posts

If you want a face on support this quarter, pre-render the answers people ask most as short Sume avatar clips, match each incoming question to one by intent, and let a live agent or chat handle the long tail. Real-time conversational avatars such as Tavus Griffin are not an option for customers yet: Tavus says Griffin-Lite is not available to customers and is open only to selected trusted testers.

Why the split works

Support questions are heavily repeated. A handful of intents (reset a password, find an invoice, change a plan) cover a large share of tickets, and each has one correct answer. A clip answers it the same way every time, can be reviewed before anyone sees it, and costs nothing per viewer after the render. The rest are specific, emotional, or account-dependent, and need a person or a conversation.

Real-time vs rendered, as of 2026-10-05
QuestionTavus Griffin-Lite (tavus.io)Sume avatar clip
AvailabilityResearch preview, selected trusted testers, not customersAvailable now through the API
InputSingle reference photograph, live audio and videoAvatar handle plus a script
Latency0.43 s average audio-to-video latency on H100sAsynchronous job: queued, processing, completed
LengthNot stated on the pageScript estimated at 4-60 seconds
PriceNot disclosedPer second by quality tier

Building the 20

Pull your last quarter of tickets and group them by intent. Write a 20-40 second script for each of the top 20, with one action per clip. Submit each as an avatar video through a preview so a reviewer approves the first frame before money is spent on the full render. Submit with an Idempotency-Key per intent and version so a retry never bills twice, and use mode: "webhook" so your server hears about completion instead of polling.

Because Sume accepts valid paid jobs as queued when the workspace is at its concurrency limit, you can submit all 20 at once; see generation admission for queue capacity on your plan. The signed completion event is described on the webhooks page.

The routing layer

Your side owns routing. A minimal version keeps a table from intent to avatar_video_id and video_url, serves the clip when the classifier is confident, and hands off when it is not.

The classifier threshold is yours to set; start strict, because a wrong clip is worse than a handoff. Always show the fallback beside the clip ("Not what you needed? Chat with us").

CLIPS = {
    "reset_password": "https://media.sume.com/artifacts/reset.mp4",
    "find_invoice": "https://media.sume.com/artifacts/invoice.mp4",
}

def route(intent, confidence, threshold=0.85):
    if intent in CLIPS and confidence >= threshold:
        return {"type": "clip", "url": CLIPS[intent],
                "fallback": "chat"}
    return {"type": "handoff", "to": "live_agent"}

print(route("reset_password", 0.93))
print(route("billing_dispute", 0.97))

What this does not do

A clip cannot listen, interrupt, see a screen, or adapt to a follow-up. It is also not a way to pass a bot off as a person, so say plainly that the presenter is an AI avatar. When a policy or price changes, the clip is stale until you re-render it, so keep the script, version, and review date next to each avatar_video_id.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume