Real-time AI avatar pricing: live minutes vs one render
Live avatars bill by conversation minute and concurrent stream; a rendered clip is paid once and watched by anyone. Tavus plan numbers and the break-even.
A real-time AI avatar is billed by the minute someone is in the conversation, with a cap on how many conversations run at once. A rendered avatar clip is billed once per render and can then be watched by any number of people. Tavus's published plans show the first model clearly; Sume's Avatar Video is the second. The break-even is how many times the same answer gets asked.
Below are Tavus's listed plan numbers (read 2026-10-03) and the arithmetic on them. Sume's own prices change with tier and are read from GET /v1/catalog, so this post does not copy a number for them.
What do live avatar plans actually charge?
Tavus's developer plans, from its pricing page: Basic is free, with 25 minutes of AI conversational video, 25 stock replicas and 1 concurrent stream. Starter is $59 a month for 100 minutes, up to 3 concurrent streams, with $0.37 per extra minute. Growth is $397 a month for 1,250 minutes, up to 10 concurrent streams, with $0.32 per extra minute and conversation recordings included. Enterprise is custom. The page also lists consumer PALs plans at $20 and $50 a month for 150 and 500 minutes of calls.
| Plan | Monthly price | Included minutes | Concurrent streams | Extra minute |
|---|---|---|---|---|
| Basic | $0 | 25 | 1 | not offered |
| Starter | $59 | 100 | Up to 3 | $0.37 |
| Growth | $397 | 1,250 | Up to 10 | $0.32 |
| Enterprise | Custom | Custom | Custom | Volume discounts |
What does the arithmetic look like?
Take Growth. Its included rate works out to about $0.32 a minute ($397 over 1,250 minutes, our division), which matches the listed overage. Now suppose 1,000 people each spend three minutes asking the same onboarding question. That is 3,000 conversation minutes, past Growth's 1,250 included ones, before you consider that only 10 can be live together. This is illustration, not a quote: your own mix of questions and viewers will differ.
A rendered clip answers the same 1,000 people with one render. The viewers press play on a public video URL; the job is not running while they watch. Sume's results can include public media.sume.com video artifacts, so there is no per-viewer session to meter (Generate avatar video).
Where does a live avatar still win on cost?
When the answers differ. If every viewer asks something new, you cannot render the answers in advance, and the per-minute model is paying for something a clip cannot do. It also wins when a conversation replaces a person's time, as in a support or tutoring call, where the comparison is a human hourly rate, not another video.
- Same answer, many viewers: render once.
- Different answer each time: pay for conversation minutes.
- Open-ended but narrow: render the top questions as clips and route the rest to a live agent.
- Concurrency is a hard ceiling on live plans; plan for the peak, not the average.
What other costs sit around each model?
On Tavus's listed plans, custom replica trainings are capped by plan (3 on Starter, 7 on Growth), extra replicas cost $65 or $40, and recordings are listed as included from Growth up (read 2026-10-03). If you want your own face in a live agent, plan for trainings as well as minutes.
On the Sume side the extra cost is rework. A script change means a new render, while preview stills let you approve framing before the full render, and the final quality tier can be set at that step (Avatar video previews). Inline captions are a separate add-on on the estimate, per the docs. Count renders, including the ones you throw away.
How do I compare the two for my case?
Write down three numbers: how many distinct answers you give, how many viewers hear each, and how long each takes. The live side costs roughly viewers times minutes, bounded by concurrent streams. The rendered side costs the count of distinct answers, times the number of rewrites. When viewers per answer is large, rendering wins by a wide margin; when it is close to one, a conversation is simpler and costs about the same.
Do not forget the free tier on each side for the test. Tavus lists 25 free minutes on Basic (read 2026-10-03), which is enough to judge how a live agent feels. For Sume, run an avatar-video preview first and judge the framing before paying for the full render.
How do I get Sume's real number?
Use the catalog: the API exposes rates by tier and the plus, standard and max quality options each have their own. Estimate one render per distinct answer, multiply by your count of answers, and compare that to the live minutes you expect. Because the clip is capped at 4-60 seconds, long answers split into several renders. For how Tavus's video-generation price compares per second, see Tavus per-minute video price vs Sume avatar per second.
Sources
Related posts
More in Pricing
- Sume plans: Pro $40, Startup $120, Scale $400 and what they limit
What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.
- TTS cost by character: the same sentence in four languages
Sume bills TTS per character, not per second. One sentence counted in English, French, German and Korean, with the 1-cent floor and the 20,000-character cap.
- Which Sume API calls are free: balance, usage, catalog, filter check
Balance, usage and catalog reads cost nothing on Sume, and neither does the video-filter check. What is billed, what is refunded, and what a 402 means.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
Written by Sume