Avatar video scene_previews: one still per scene, one shared set
Multi-scene avatar previews return scene_previews, one still per scene. Later stills continue scene 0's pose, so review scene 0 first.
When you preview a multi-scene avatar video, Sume returns preview_image_url for the primary still (scene 0) and scene_previews[] with one still per input scene. For a shared scene, later stills are pose-anchored continuations of the first frame, so you judge the set as one look rather than as independent images.
That shapes how you review: scene 0 sets the pose, framing and background that every later still builds on.
What does the preview resource return?
POST /v1/avatar-video-previews creates the preview with the same body as an avatar video: exactly one of script or video_inputs, plus optional product_image, scene, quality, aspect_ratio, title and captions. The response includes job polling URLs and an avatar_video_preview_id.
When ready, GET /v1/avatar-video-previews/{id} exposes preview_image_url, scene_previews[] when available, and resource_status and job_status. Prefer those two over the legacy status field: resource_status answers whether the preview is ready, job_status answers whether the job is still running.
Why does scene 0 matter so much?
Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. In a preview of that shape the first still fixes the avatar pose and framing, and later stills continue from it. If scene 0 is wrong, the later stills inherit the problem, so fix it before you judge anything else.
So check scene 0 for the things you cannot change after approval: avatar, background, aspect ratio and product placement. These come from avatar_handle, scene and aspect_ratio, which are structural fields.
Which review steps catch real problems?
Work through the stills in order and keep the checks short.
- Scene 0: is the avatar the one you intended, and is the framing right for your aspect ratio (default 9:16)?
- Product shots: does the
product_imageread clearly, or is it covered by the avatar? - Later scenes: do they feel like the same room and pose, or did the continuation drift?
- Captions: they are stored on create and are never burned into stills, so do not judge them here.
What do I do when a still is wrong?
If only the look is off, call POST /v1/avatar-video-previews/{id}/regenerate. It reuses the stored request and refreshes only the first-frame stills, returning the same avatar_video_preview_id with a new preview-only job. If the script, video_inputs, avatar, scene or aspect ratio is wrong, create a new preview.
When it looks right, call generate-video on the preview id. Sume reuses the preview first frame when available. An empty body keeps the quality chosen at create; an optional quality overrides only the final render tier.
curl -X POST https://api.sume.com/v1/avatar-video-previews/avp_123/generate-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-preview-avp_123-final" \
-d '{"quality":"max"}'Is anything in the preview reused for free?
Stills are tier-independent and always reused, so approving once and then moving between standard, plus and max does not need a new preview. Admission, pre-spend, reservation, provider submit and readback all use the effective tier you pass at generate-video.
Read the exact schemas in the live OpenAPI reference linked from the Avatar video previews page. See also regenerate stills or create a new preview.
Sources
Related posts
More in Sume Avatar 1.0
- Cancel an avatar video job: the 409 job_generation_already_started
You can cancel an avatar video job only before generation starts. After that Sume returns 409 job_generation_already_started and the job runs to completion.
- Create an AI avatar from profile traits: the props input
Avatar 1.0 can build a reusable avatar from structured traits, not a prompt or photo. The props input takes ethnicity, sex and age. When to use it.
- Create an AI avatar from a reference image: URL rules and cost
Sume turns a public HTTPS photo into a reusable avatar for $0.95. The request, the URL checks that reject a bad image, and how to use the handle in videos.
- Face swap video_url rejected: signed and private URLs explained
Sume face swap needs a public HTTPS video_url. Signed or private URLs, localhost and provider task URLs are rejected before generation. How to host the clip.
Written by Sume