AI avatars for a 5-person team: 5 avatars, 4 clips each, budget

Five team avatars and four 20-second clips each cost $4.75 to create plus $98.00 of Plus video: $102.75 at Sume's listed rates, with Standard and Max compared.

5 min readSume
All posts

Giving five people on a team their own avatar and four 20-second clips each costs $4.75 for the avatars and $98.00 of Plus video, $102.75 in total at Sume's listed rates (read 2026-10-03). That is 400 seconds of video across 20 clips. On Standard the video is $73.60; on Max it is $220.00.

The question usually comes from a sales or enablement lead who wants every presenter to appear in their own recognisable clips without booking a studio. The cost model is two separate meters: one flat charge per avatar, and a per-second charge per video.

Two meters, not one

Creating an avatar is a one-time charge for each avatar you create, from a prompt, structured profile traits, or a reference image, as described in Create new avatar. Every avatar gets a handle that you pass as avatar_handle when you generate videos with the avatar video route.

Because the avatar is reusable, the avatar charge does not grow with the number of clips. Five presenters cost five creation charges whether each makes one clip or fifty. That is why the fixed part of this budget is small and the per-second part dominates.

Team budget at list rates, read 2026-10-03
Line itemCost
Create 5 avatars$4.75
20 clips x 20 s on Standard$73.60
20 clips x 20 s on Plus$98.00
20 clips x 20 s on Max$220.00

Totals by tier

Adding the avatar charge to each video line gives the complete first-month budget:

Complete budget, read 2026-10-03
TierAll 5 avatars and 20 clips
Standard$78.35
Plus$102.75
Max$224.75

Check the numbers yourself

The calculation is a few lines of Python. Replace the constants if the rates change; the pricing page is the source of truth.

import asyncio

PLUS = 0.245        # USD per second, no product image
CREATE = 0.95       # USD per avatar, once

async def main() -> None:
    presenters, clips_each, seconds = 5, 4, 20
    avatars = presenters * CREATE
    video = presenters * clips_each * seconds * PLUS
    print(f"avatars ${avatars:.2f}, video ${video:.2f}, total ${avatars + video:.2f}")

asyncio.run(main())

What this does not include

  • Product images: attaching one raises the per-second rate a little, so a product-demo clip costs slightly more than the figures above.
  • Retakes: any clip you regenerate is another paid job.
  • Consent: if an avatar is based on a real person's photo, get that person's agreement first. Sume cannot decide that for you.
  • Concurrency: five people queueing clips at once share the workspace's generation concurrency and queue, described in Generation admission.

Planning the roster

Before spending anything, decide how each presenter's avatar will be made. The docs offer three inputs: a text prompt, structured profile traits (called props in the API, for ethnicity, sex and age), or a reference image URL. A prompt or profile avatar is a fictional presenter; an image avatar is based on the picture you supply, which should be a public HTTPS URL. Giving each team member a distinct, memorable handle such as sales_maya or support_dev makes the later avatar_handle fields self-explanatory, and Sume stores the handle without a leading @.

If these avatars depict real colleagues, write down who agreed. A shared one-page consent note costs nothing and avoids awkward conversations later. Sume does not offer legal advice, so treat the tool as a producer of clips and keep your own records of permission.

Spreading the spend across a month

Twenty 20-second clips do not need to be rendered on one day. Because the workspace has a generation concurrency limit set by plan and a queue behind it, a burst of 20 submissions is accepted as long as queue capacity remains, and the jobs wait in queued until a slot opens. On a plan with a small queue, submit in waves of a few clips and let each wave finish. The admission docs list the limits by plan, and the safe approach is to read your effective limit rather than assume.

A second reason to stagger is learning. The first clip from each presenter tells you whether the script length, tone and background suit the person. Fix the template after clip one, not after clip twenty.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume