AI avatars for a 5-person team: 5 avatars, 4 clips each, budget
Five team avatars and four 20-second clips each cost $4.75 to create plus $98.00 of Plus video: $102.75 at Sume's listed rates, with Standard and Max compared.
Giving five people on a team their own avatar and four 20-second clips each costs $4.75 for the avatars and $98.00 of Plus video, $102.75 in total at Sume's listed rates (read 2026-10-03). That is 400 seconds of video across 20 clips. On Standard the video is $73.60; on Max it is $220.00.
The question usually comes from a sales or enablement lead who wants every presenter to appear in their own recognisable clips without booking a studio. The cost model is two separate meters: one flat charge per avatar, and a per-second charge per video.
Two meters, not one
Creating an avatar is a one-time charge for each avatar you create, from a prompt, structured profile traits, or a reference image, as described in Create new avatar. Every avatar gets a handle that you pass as avatar_handle when you generate videos with the avatar video route.
Because the avatar is reusable, the avatar charge does not grow with the number of clips. Five presenters cost five creation charges whether each makes one clip or fifty. That is why the fixed part of this budget is small and the per-second part dominates.
| Line item | Cost |
|---|---|
| Create 5 avatars | $4.75 |
| 20 clips x 20 s on Standard | $73.60 |
| 20 clips x 20 s on Plus | $98.00 |
| 20 clips x 20 s on Max | $220.00 |
Totals by tier
Adding the avatar charge to each video line gives the complete first-month budget:
| Tier | All 5 avatars and 20 clips |
|---|---|
| Standard | $78.35 |
| Plus | $102.75 |
| Max | $224.75 |
Check the numbers yourself
The calculation is a few lines of Python. Replace the constants if the rates change; the pricing page is the source of truth.
import asyncio
PLUS = 0.245 # USD per second, no product image
CREATE = 0.95 # USD per avatar, once
async def main() -> None:
presenters, clips_each, seconds = 5, 4, 20
avatars = presenters * CREATE
video = presenters * clips_each * seconds * PLUS
print(f"avatars ${avatars:.2f}, video ${video:.2f}, total ${avatars + video:.2f}")
asyncio.run(main())
What this does not include
- Product images: attaching one raises the per-second rate a little, so a product-demo clip costs slightly more than the figures above.
- Retakes: any clip you regenerate is another paid job.
- Consent: if an avatar is based on a real person's photo, get that person's agreement first. Sume cannot decide that for you.
- Concurrency: five people queueing clips at once share the workspace's generation concurrency and queue, described in Generation admission.
Planning the roster
Before spending anything, decide how each presenter's avatar will be made. The docs offer three inputs: a text prompt, structured profile traits (called props in the API, for ethnicity, sex and age), or a reference image URL. A prompt or profile avatar is a fictional presenter; an image avatar is based on the picture you supply, which should be a public HTTPS URL. Giving each team member a distinct, memorable handle such as sales_maya or support_dev makes the later avatar_handle fields self-explanatory, and Sume stores the handle without a leading @.
If these avatars depict real colleagues, write down who agreed. A shared one-page consent note costs nothing and avoids awkward conversations later. Sume does not offer legal advice, so treat the tool as a producer of clips and keep your own records of permission.
Spreading the spend across a month
Twenty 20-second clips do not need to be rendered on one day. Because the workspace has a generation concurrency limit set by plan and a queue behind it, a burst of 20 submissions is accepted as long as queue capacity remains, and the jobs wait in queued until a slot opens. On a plan with a small queue, submit in waves of a few clips and let each wave finish. The admission docs list the limits by plan, and the safe approach is to read your effective limit rather than assume.
A second reason to stagger is learning. The first clip from each presenter tells you whether the script length, tone and background suit the person. Fix the template after clip one, not after clip twenty.
Sources
Related posts
More in Pricing
- AI image cost per image: Gemini, Luma, Ideogram and GPT Image lists
Vendor lists per image today: Gemini Nano Banana 2 $0.067 (1K), Luma Uni-1.1 $0.0404 (2K), Ideogram 4.5 $0.03-$0.22, GPT Image 2.5 $0.09366 (xhigh, 1024).
- Which AI image models cost under 4 cents on the Sume image API?
Six Sume image models list under $0.04 per image: Soul, Grok Imagine, Qwen Image, Imagen 4 Fast, Seedream 4.0 and Flux 2 Pro. What each takes, for drafts.
- AI lip sync API cost per second: H3 Max 480p, 768p, 1080p vs Fabric
On Sume, H3 Max lip-sync is $0.0625, $0.10 or $0.20 per audio second by resolution; VEED Fabric is $0.10 or $0.1875. A 14.8 s clip costs $0.94 to $3.00.
- AI music for 100 short videos: $12.50 flat on Sume Music
One hundred Sume Music generations cost $12.50 at the fixed $0.125 per accepted generation, whatever the prompt length. What that does and doesn't cover.
Written by Sume