AI avatar sales agent: live SDR or personalised clips?
Tavus builds live SDR avatars on a per-minute plan. Sume renders one 4-60 second avatar clip per lead from a script. How to pick, plus a Python loop.
An AI avatar sales agent comes in two shapes. A live one holds a video conversation with a prospect; Tavus's documentation lists a Sales Development Rep PAL as one of four example use cases, next to an interviewer, customer support and medical intake (read 2026-10-03). A rendered one is a short personalised clip sent to each lead. Sume does the second: one script per lead in, one avatar video of 4-60 seconds out. It does not take live calls.
Pick live when the lead needs to ask questions in the moment, and rendered when the job is to get a clear, reviewed first message in front of many people.
When does a live avatar SDR make sense?
When the next sentence depends on the prospect. Tavus's PAL pipeline supports knowledge-base documents (document_ids), memory tagged per participant and a custom_greeting, according to its create-conversation reference (read 2026-10-03). That fits qualification, pricing questions and objection handling.
The cost model follows the conversation. Tavus lists plans from a free tier of 25 minutes with one concurrent stream, to Starter at $59 a month with 100 minutes and up to 3 concurrent streams, and Growth at $397 with 1,250 minutes and up to 10 (Tavus pricing, read 2026-10-03). The limit that matters for outbound is concurrency: each live prospect occupies a stream while they talk.
When do personalised clips beat a live agent?
For the cold first touch. A clip is watched on the prospect's schedule, needs no concurrency, and can be reviewed before it goes out. With Sume you can render a preview first and approve the first frames before paying for the full render (Avatar video previews).
Personalisation is whatever you put in the script. Sume does not look anyone up; your code fills the name, company or product line into the text and submits. Give each submit its own Idempotency-Key, derived from the lead id, so a retry does not create a second clip.
import os, requests
API = "https://api.sume.com/v1/avatar-1.0/talking-video"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
leads = [
{"id": "001", "first": "Dana", "product": "the onboarding kit"},
{"id": "002", "first": "Sam", "product": "the analytics add-on"},
]
for lead in leads:
script = (f"Hi {lead['first']}, I recorded this short note about "
f"{lead['product']}. Reply if you want a walkthrough.")
r = requests.post(
API,
headers={**HEADERS, "Idempotency-Key": f"outreach-{lead['id']}"},
json={"avatar_handle": "sume_clawra", "script": script,
"aspect_ratio": "9:16", "quality": "standard",
"mode": "webhook",
"webhook_url": "https://example.com/hooks/sume"},
timeout=30,
)
r.raise_for_status()
print(lead["id"], r.json()["data"]["status_url"])How do the two compare for outbound?
The deciding fields are who has to be present and what you review.
| Question | Live avatar (Tavus) | Rendered clip (Sume) |
|---|---|---|
| Can the lead ask a question | Yes | No |
| Lead must be present while it runs | Yes | No, they watch later |
| What limits volume | Concurrent streams per plan | Your workspace's job queue |
| Can you review exactly what is said first | Not in advance | Yes, script and preview stills |
| Unit of cost | Minutes of conversation | One render per clip |
Should I preview every clip?
Preview the template, not every lead. When the only difference between clips is a name or a product line, create one avatar-video preview with a sample lead, check preview_image_url, and then submit the batch. The preview stills are tier-independent, and generate-video can override quality for the final render only, so approving a draft framing does not lock you into a tier (Avatar video previews).
Spot-check a few rendered outputs from the batch afterwards. Names are the usual failure: a script that reads a handle or a company in an odd way is cheaper to fix in the template than in a hundred sent emails.
What about the follow-up call?
A clip works best as the opener to a human or live-agent conversation, not a replacement for it. The clip states one useful thing, and the next step is a reply, a booking link or a live session. Tavus's conversation API can start a PAL from a link (conversation_url) or send it into a meeting with meeting_url (create-conversation reference, read 2026-10-03), so a rendered clip followed by a live link is a workable sequence.
Keep records of what was said. With a rendered clip, the exact script is the record; store it with the lead id and the Idempotency-Key. That is a small but real advantage when someone later asks what a prospect was told.
What should the clip say?
Keep it under the length of the script window and make the ask one step: reply, book, or watch the product. If the lead replies with a question, that is the moment to hand over to a person or to a live agent. The webhook posts a job.completed event with the artifact URL; the verifier on your side should refuse to run with an empty secret (Webhooks).
- Use a clip for the first touch and a live agent for the reply.
- Derive the Idempotency-Key from the lead id.
- Do not claim the avatar is a person; say it is an AI-generated presenter.
- Read the preview before sending a batch.
Sources
Related posts
More in Use cases
- AI backing track generator: an instrumental to sing or play over
Generate an AI backing track by prompt: tempo, key, form, no vocals. Sume returns one mixed MP3, no stems or click track, so trim and check the take by ear.
- AI character series on Shorts: avoid the same situation each time
YouTube's inauthentic content policy flags characters in identical situations with the same outcomes. How to keep an AI character and vary the story in Sume.
- AI classroom background music for lesson videos, under narration
Make calm instrumental music for a lesson video: a prompt that keeps vocals out, a Python script, and a Timeline bed that ducks under the teacher's voice.
- AI person in a Meta ad: is the label next to Sponsored?
Meta puts AI info next to Sponsored when its own tools make a photorealistic human. For outside tools it describes About this ad. How to add your own cue.
Written by Sume