Preview an AI avatar video before paying for the full render

Sume avatar video previews make first-frame stills only. Approve them, then call generate-video; you can change the final quality tier without a new preview.

5 min readSume
All posts

On Sume you can approve an avatar video's first frames before the full render starts. POST /v1/avatar-video-previews generates the first-frame still stage only and does not start the talking-video render. When the stills look right, call generate-video on the preview id, and the final render reuses those stills.

The four routes

From Avatar video previews, read 2026-10-08:

Preview routes, as of 2026-10-08
RoutePurpose
POST /v1/avatar-video-previewsCreate the preview job
GET /v1/avatar-video-previews/:idRead the preview resource and stills
POST /v1/avatar-video-previews/:id/regenerateRefresh only the first-frame stills from the stored request
POST /v1/avatar-video-previews/:id/generate-videoStart the normal Avatar Video workflow

What you can change after approval

This is the useful part. Preview stills are tier-independent, and Sume always reuses them. An empty body (or {}) on generate-video keeps the quality tier you chose at preview create. Send quality to override only the final render tier. So if you approve a preview at the default plus tier and later decide to run the final at standard or max, you do not need a new preview. Admission, pre-spend, ledger reservation, provider submit and readback all use the effective tier.

Structural fields are different. A change to script, video_inputs, avatar_handle, scene or aspect_ratio needs a new preview.

  • Can change on generate-video: quality.
  • Needs a new preview: script, video_inputs, avatar_handle, scene, aspect_ratio.
  • Captions: stored at preview create and applied at generate-video; never burned into stills.

When a preview is worth it

Use a preview for multi-scene video_inputs where you want one still per scene, for product-in-hand shots where composition is the risk, and for any batch where one bad framing would be repeated across many renders. For a plain talking head with a script you have used before, a direct render through Generate avatar video is simpler.

A ready preview exposes preview_image_url as the primary still (scene 0 for multi-scene) and scene_previews[] with one still per scene when available. Read resource_status for readiness and job_status for the job, rather than the legacy status field.

Same limits as the full video

The duration window is unchanged: an estimated 4-60 seconds inclusive. Media inputs are public HTTPS URLs. If a still is wrong, call regenerate rather than rewriting the request. It returns the same avatar_video_preview_id and a new preview-only job.

A cautious workflow

Create the preview, wait for resource_status to show ready, and open preview_image_url and each entry in scene_previews. Look for the face, the product in hand, the framing at your chosen ratio, and any background artifacts. If one still is off, call regenerate and recheck. When the stills are right, call generate-video with an empty body to keep your tier, or send quality to change only the final render tier.

Keep notes. A preview id belongs to one request, so store it next to the script version it came from. If you edit the script, make a new preview instead of reusing the old id, because structural changes are not applied through generate-video.

If you use captions, set them when you create the preview. They are stored and applied at generate-video time, never burned into the stills, so previews will not show them.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume