Avatar video quality tiers: test one script at all three
How to choose between standard, plus and max in Sume Avatar 1.0: render one script three times, compare on a checklist, and keep the cheapest tier that passes.
Pick the quality tier by testing, not by guessing. Render one representative script at standard, plus and max, review the three files side by side, and ship the cheapest tier that passes your checklist. Sume Avatar 1.0 accepts those three values for quality, and plus is the default when you omit it.
What the tiers are
The docs describe the tiers by behavior, not by a fixed feature list. Use this table as your starting hypothesis, and let the test confirm it.
| Value | Documented behavior | Try it first when |
|---|---|---|
| standard | Fastest Sume execution path | Drafts, internal review, high volume |
| plus | Default. Balanced quality path | Most customer-facing clips |
| max | Highest quality tier, slower turnaround | A hero clip on the landing page |
Run the test
Use a script that stresses what you care about: a name that is easy to mispronounce, a number, a pause, and a close-up phrase. Keep it inside the accepted 4 to 60 second window so the three renders are comparable.
Send the same body three times, changing only quality. Give each request its own idempotency key, such as tier-test-standard, so a retry of one tier never creates a duplicate and the three tiers stay distinct jobs. Poll each job until it is completed, then fetch the result.
Score them on a checklist
Watch all three on the device your audience uses, with the sound on. Score each file pass or fail on the same items.
- Lip movement matches the words, including the hard name.
- Face and skin look stable between sentences.
- Pauses and pacing feel natural at phone size.
- The clip still reads well after your captions are added.
- Turnaround fits your publishing schedule.
Save the cost of a bad full render
If the problem is composition rather than quality, do not re-render the whole clip. An avatar video preview generates only the first-frame stills, so you can approve the framing before you pay for a full talking video. Then call generate-video on the preview id.
Read each tier's current price from your own usage and the pricing pages rather than from this post, because this article does not quote prices.
Decide and record it
Write the winning tier down next to the script type, for example plus for product clips and standard for drafts. Re-run the three-way test when you change the avatar, the aspect ratio or the script style. A tier choice that was right for a 15-second hook may not hold for a 55-second explainer.
Sources
Related posts
More in Sume Avatar 1.0
- Griffin-style follow-ups with rendered avatar clips and branching
Griffin-Lite reacts live. Until you can use it, approximate a guided conversation with a set of pre-rendered Sume avatar clips and your own branching logic.
- Interactive avatar or avatar video? A five-question test
Griffin-Lite is a research preview. Five questions tell you whether you need a live avatar or a rendered Sume Avatar 1.0 clip, and what to ship this quarter.
- Name your avatars: a handle scheme that fits 2 to 30 characters
Sume avatar handles allow letters, digits, periods and underscores, 2 to 30 characters. Build a team naming scheme that passes validation and stays readable.
- Let users pick an avatar in your app with GET /v1/avatar-1.0/avatars
Build an avatar picker on the Sume Avatar 1.0 list route: read ready avatars server-side, cache them, and pass the chosen handle to talking-video.
Written by Sume