Avatar first-frame preview checklist: seven checks before final render
Before paying for the full avatar render, check seven things on the preview stills: handle, framing, background, product, ratio, scenes, caption intent.
A first-frame preview is the cheap place to catch mistakes: it generates stills, not the talking video, and you only start the full render when you approve. Check seven things on the stills, in this order, and call regenerate or create a new preview depending on what you find.
The seven checks
- Right avatar: the person in the still is the handle you meant. Handles from another workspace return 404, so a wrong face usually means a wrong handle in your own table.
- Framing: head and shoulders are not cropped awkwardly for the aspect ratio you chose. Aspect ratio is
1:1,3:4,9:16,4:3or16:9, default9:16. - Background: matches your scene prompt or scene photo, and nothing in it contradicts the script (a kitchen for a legal explainer).
- Product: if you sent a
product_image, it is visible and not hidden behind a hand. - Scene continuity: for multi-scene
video_inputs, later stills are pose-anchored continuations of the first frame; check each one inscene_previews[]. - Safe areas: the lower third of a 9:16 frame is where captions land, so nothing important should sit there.
- Captions intent: stored at preview create but never burned into the stills, so do not expect to see them.
What each fix costs
The docs separate what a preview can absorb from what it cannot. regenerate reuses the stored request and refreshes only the first-frame stills, returning the same avatar_video_preview_id with a new preview-only job. Changes to structural fields (script, video_inputs, avatar_handle, scene, aspect ratio) need a new preview.
| Problem | Action |
|---|---|
| Still looks off but the request is right | Call regenerate on the same preview |
| Wrong script or scene | Create a new preview |
| Wrong crop | Change aspect ratio, new preview |
| Want a different tier | No new preview; pass quality at generate-video |
Then generate
When the stills are right, call generate-video on the preview id. An empty body keeps the quality chosen at create; an explicit quality overrides it for the final render only. Sume reuses the approved first frame, and the captions you stored at create apply at this step. Poll resource_status for readiness and job_status for the job, as the status post explains.
The preview must still fit the 4 to 60 second window, since it uses the same estimate as the final video. Check your script length before you create it, and split anything longer.
Make it a team habit
Have the person who owns the brand look at the stills, not only the person who wrote the script. A reviewer who sees the first frame spots a wrong background in seconds, and a regenerate is far less costly than a rendered clip that has to be thrown away.
Why stills are not the whole story
A first frame cannot show motion, lip sync or voice. It shows composition. Pair the preview with a short, cheap standard-tier test render when the script is new or the voice matters, and keep max for the final. Because preview stills do not depend on the tier, you can approve a still and then pick any tier at generate-video.
Sources
Related posts
More in Sume Avatar 1.0
- 10 products, one 18-second avatar clip each: cost with a product image
Ten 18-second Avatar Video clips cost $33.12 at standard, $44.10 at plus and $99.00 at max; a product image adds $1.80 on standard.
- What a scripted AI avatar clip cannot do: seven limits and fixes
A Sume avatar clip cannot take questions, run past 60 seconds or switch presenter mid-video. Seven documented limits, each with a workaround.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume