Pre-rendered avatar greetings per visitor segment, not a live avatar
Instead of a live avatar for each visitor, render one short Sume avatar clip per segment ahead of time. Per-tier cost for six 12-second greetings.
If your goal is a friendly face on a website that says something relevant to each visitor, you can often get there with a handful of pre-rendered clips, one per segment, rather than a live avatar that talks to each person. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session. That is a good fit when you know the segments in advance: new visitor, returning customer, pricing-page visitor, trial user, partner, and so on.
This post gives the cost for six 12-second greetings at each Avatar 1.0 quality tier and says plainly where a pre-rendered approach stops working.
The cost of six greetings
Avatar Video bills per second at a public rate per tier. The 4-second floor and 60-second ceiling apply to each video. Six clips of 12 seconds each is 72 seconds of video in total. The rates below are the no-product rates from packages/provider-pricing (sume_avatar_video_fixed_tier_rates_2026_07_02).
| Tier | Rate per second | One 12 s clip | Six clips (72 s) |
|---|---|---|---|
| standard | $0.184 | $2.208 | $13.248 |
| plus | $0.245 | $2.94 | $17.64 |
| max | $0.55 | $6.60 | $39.60 |
Why it works
One clip serves every visitor in a segment, so cost scales with the number of segments and not with the number of viewers. Add a new segment, render one more clip. Change the offer, re-render only the segments that mention it. And because the avatar is a reusable handle that you created once for $0.95, every greeting uses the same face and voice.
Where it stops working
A pre-rendered clip cannot answer a question the visitor just typed. It cannot react to what they say, and it cannot show their name unless you render a clip per name. If you need that, you are describing a live conversational product, which is outside what Sume avatars do. A reasonable compromise is a live chat or form for the question, with the avatar clip as the greeting and the recap.
Render one segment
Submit a talking video with the segment script, then poll the job. Use standard while you review copy and move to plus or max for the final cut. The preview route lets you approve first-frame stills before paying for the full render.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: greeting-returning-001" \
-d '{
"avatar_handle": "product_host",
"script": "Welcome back. Your dashboard has two new reports waiting.",
"quality": "standard",
"aspect_ratio": "16:9"
}'Choosing the segments
Start from the questions your visitors already answer without typing: which page they are on, whether they are signed in, which plan they have, which campaign link they came from. Each of those can map to a segment. Keep the first version small. Three to six segments cover most sites, and each extra segment is another clip you must keep current when your product changes.
Write each greeting as one idea and aim for 10 to 15 seconds. A script that is too short to reach the 4-second minimum can be padded with a silence beat, and a script that grows past 60 seconds must be split into two jobs. A greeting is not the place for either.
Keeping greetings fresh
Treat the clips like any other piece of site content. Store the script, the avatar handle, the quality tier and the job id next to each segment. When the copy changes, render a new clip and swap the URL, and keep the old URL until the new one is live. Since each render is a job with a reservation, you can add a spend cap in your own tooling and know the maximum for a refresh: for six 12-second clips at plus, that is $17.64 for the whole set.
Finally, tell visitors that the person on screen is synthetic. A one-line label beside the player is enough and it avoids surprises.
- Each clip is 4 to 60 seconds, and the greeting fits in 10 to 15.
- The avatar handle is the same in every segment script.
- Captions are on if the clip plays muted, using the
captionsfield on the request. - Each job id is stored beside its segment so you can fetch the result again later.
Sources
Related posts
More in Use cases
- Product photo to a 6-second vertical ad: a still is a static hold
A still in a Timeline video slot is held, not animated: motion is accepted but ignored with a motion_ignored warning. Default 1080x1920, $0.10 for 6 seconds.
- Remove a date stamp from a scanned photo: crop first, then a mask edit
Remove an orange date stamp from a scanned family photo: crop it for free in Pillow, or mask it and edit with openai/gpt-image-2.5 from $0.0094 an image.
- Remove dust, lint and wrinkles from a product photo with an AI edit
Clean dust, lint and fabric wrinkles from a product photo with ideogram/ideogram-v4.5: one reference, no mask, $0.0375 to $0.275 per image on Sume.
- Remove the photographer's reflection from a window photo: masked edit
Remove your own reflection from a shop window photo: mask it and send a gpt-image-2.5 edit with mask_url, $0.0094 to $0.0835 per image on Sume. Checks inside.
Written by Sume