AI avatar sales agent: live SDR or personalised clips?

Tavus builds live SDR avatars on a per-minute plan. Sume renders one 4-60 second avatar clip per lead from a script. How to pick, plus a Python loop.

5 min readSume
All posts

An AI avatar sales agent comes in two shapes. A live one holds a video conversation with a prospect; Tavus's documentation lists a Sales Development Rep PAL as one of four example use cases, next to an interviewer, customer support and medical intake (read 2026-10-03). A rendered one is a short personalised clip sent to each lead. Sume does the second: one script per lead in, one avatar video of 4-60 seconds out. It does not take live calls.

Pick live when the lead needs to ask questions in the moment, and rendered when the job is to get a clear, reviewed first message in front of many people.

When does a live avatar SDR make sense?

When the next sentence depends on the prospect. Tavus's PAL pipeline supports knowledge-base documents (document_ids), memory tagged per participant and a custom_greeting, according to its create-conversation reference (read 2026-10-03). That fits qualification, pricing questions and objection handling.

The cost model follows the conversation. Tavus lists plans from a free tier of 25 minutes with one concurrent stream, to Starter at $59 a month with 100 minutes and up to 3 concurrent streams, and Growth at $397 with 1,250 minutes and up to 10 (Tavus pricing, read 2026-10-03). The limit that matters for outbound is concurrency: each live prospect occupies a stream while they talk.

When do personalised clips beat a live agent?

For the cold first touch. A clip is watched on the prospect's schedule, needs no concurrency, and can be reviewed before it goes out. With Sume you can render a preview first and approve the first frames before paying for the full render (Avatar video previews).

Personalisation is whatever you put in the script. Sume does not look anyone up; your code fills the name, company or product line into the text and submits. Give each submit its own Idempotency-Key, derived from the lead id, so a retry does not create a second clip.

import os, requests

API = "https://api.sume.com/v1/avatar-1.0/talking-video"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
leads = [
    {"id": "001", "first": "Dana", "product": "the onboarding kit"},
    {"id": "002", "first": "Sam", "product": "the analytics add-on"},
]
for lead in leads:
    script = (f"Hi {lead['first']}, I recorded this short note about "
              f"{lead['product']}. Reply if you want a walkthrough.")
    r = requests.post(
        API,
        headers={**HEADERS, "Idempotency-Key": f"outreach-{lead['id']}"},
        json={"avatar_handle": "sume_clawra", "script": script,
              "aspect_ratio": "9:16", "quality": "standard",
              "mode": "webhook",
              "webhook_url": "https://example.com/hooks/sume"},
        timeout=30,
    )
    r.raise_for_status()
    print(lead["id"], r.json()["data"]["status_url"])

How do the two compare for outbound?

The deciding fields are who has to be present and what you review.

Live SDR avatar vs rendered clip, read 2026-10-03
QuestionLive avatar (Tavus)Rendered clip (Sume)
Can the lead ask a questionYesNo
Lead must be present while it runsYesNo, they watch later
What limits volumeConcurrent streams per planYour workspace's job queue
Can you review exactly what is said firstNot in advanceYes, script and preview stills
Unit of costMinutes of conversationOne render per clip

Should I preview every clip?

Preview the template, not every lead. When the only difference between clips is a name or a product line, create one avatar-video preview with a sample lead, check preview_image_url, and then submit the batch. The preview stills are tier-independent, and generate-video can override quality for the final render only, so approving a draft framing does not lock you into a tier (Avatar video previews).

Spot-check a few rendered outputs from the batch afterwards. Names are the usual failure: a script that reads a handle or a company in an odd way is cheaper to fix in the template than in a hundred sent emails.

What about the follow-up call?

A clip works best as the opener to a human or live-agent conversation, not a replacement for it. The clip states one useful thing, and the next step is a reply, a booking link or a live session. Tavus's conversation API can start a PAL from a link (conversation_url) or send it into a meeting with meeting_url (create-conversation reference, read 2026-10-03), so a rendered clip followed by a live link is a workable sequence.

Keep records of what was said. With a rendered clip, the exact script is the record; store it with the lead id and the Idempotency-Key. That is a small but real advantage when someone later asks what a prospect was told.

What should the clip say?

Keep it under the length of the script window and make the ask one step: reply, book, or watch the product. If the lead replies with a question, that is the moment to hand over to a person or to a live agent. The webhook posts a job.completed event with the artifact URL; the verifier on your side should refuse to run with an empty secret (Webhooks).

  • Use a clip for the first touch and a live agent for the reply.
  • Derive the Idempotency-Key from the lead id.
  • Do not claim the avatar is a person; say it is an AI-generated presenter.
  • Read the preview before sending a batch.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume