Preview an AI avatar video before paying for the full render
Sume avatar video previews make first-frame stills only. Approve them, then call generate-video; you can change the final quality tier without a new preview.
On Sume you can approve an avatar video's first frames before the full render starts. POST /v1/avatar-video-previews generates the first-frame still stage only and does not start the talking-video render. When the stills look right, call generate-video on the preview id, and the final render reuses those stills.
The four routes
From Avatar video previews, read 2026-10-08:
| Route | Purpose |
|---|---|
| POST /v1/avatar-video-previews | Create the preview job |
| GET /v1/avatar-video-previews/:id | Read the preview resource and stills |
| POST /v1/avatar-video-previews/:id/regenerate | Refresh only the first-frame stills from the stored request |
| POST /v1/avatar-video-previews/:id/generate-video | Start the normal Avatar Video workflow |
What you can change after approval
This is the useful part. Preview stills are tier-independent, and Sume always reuses them. An empty body (or {}) on generate-video keeps the quality tier you chose at preview create. Send quality to override only the final render tier. So if you approve a preview at the default plus tier and later decide to run the final at standard or max, you do not need a new preview. Admission, pre-spend, ledger reservation, provider submit and readback all use the effective tier.
Structural fields are different. A change to script, video_inputs, avatar_handle, scene or aspect_ratio needs a new preview.
- Can change on generate-video: quality.
- Needs a new preview: script, video_inputs, avatar_handle, scene, aspect_ratio.
- Captions: stored at preview create and applied at generate-video; never burned into stills.
When a preview is worth it
Use a preview for multi-scene video_inputs where you want one still per scene, for product-in-hand shots where composition is the risk, and for any batch where one bad framing would be repeated across many renders. For a plain talking head with a script you have used before, a direct render through Generate avatar video is simpler.
A ready preview exposes preview_image_url as the primary still (scene 0 for multi-scene) and scene_previews[] with one still per scene when available. Read resource_status for readiness and job_status for the job, rather than the legacy status field.
Same limits as the full video
The duration window is unchanged: an estimated 4-60 seconds inclusive. Media inputs are public HTTPS URLs. If a still is wrong, call regenerate rather than rewriting the request. It returns the same avatar_video_preview_id and a new preview-only job.
A cautious workflow
Create the preview, wait for resource_status to show ready, and open preview_image_url and each entry in scene_previews. Look for the face, the product in hand, the framing at your chosen ratio, and any background artifacts. If one still is off, call regenerate and recheck. When the stills are right, call generate-video with an empty body to keep your tier, or send quality to change only the final render tier.
Keep notes. A preview id belongs to one request, so store it next to the script version it came from. If you edit the script, make a new preview instead of reusing the old id, because structural changes are not applied through generate-video.
If you use captions, set them when you create the preview. They are stored and applied at generate-video time, never burned into the stills, so previews will not show them.
Sources
Related posts
More in Sume Avatar 1.0
- Put the AI disclosure in scene one: Sume avatar video_inputs recipe
A copyable Sume Avatar 1.0 request with a 3-second disclosure scene and a 12-second message. The 15 seconds cost $2.76 at standard, $3.68 at plus.
- Video call avatar or scripted talking video: which do you need?
A real-time avatar answers people live; a scripted talking video is a file you render and review first. How to choose, and what Sume Avatar 1.0 covers.
- Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?
Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.
- Reuse one AI spokesperson across videos with an avatar handle
A Sume avatar_handle is a stable name for a ready avatar. Sume strips a leading @, so you create once and call the handle in every talking-video request.
Written by Sume