Avatar video scene_previews: one still per scene, one shared set

Multi-scene avatar previews return scene_previews, one still per scene. Later stills continue scene 0's pose, so review scene 0 first.

4 min readSume
All posts

When you preview a multi-scene avatar video, Sume returns preview_image_url for the primary still (scene 0) and scene_previews[] with one still per input scene. For a shared scene, later stills are pose-anchored continuations of the first frame, so you judge the set as one look rather than as independent images.

That shapes how you review: scene 0 sets the pose, framing and background that every later still builds on.

What does the preview resource return?

POST /v1/avatar-video-previews creates the preview with the same body as an avatar video: exactly one of script or video_inputs, plus optional product_image, scene, quality, aspect_ratio, title and captions. The response includes job polling URLs and an avatar_video_preview_id.

When ready, GET /v1/avatar-video-previews/{id} exposes preview_image_url, scene_previews[] when available, and resource_status and job_status. Prefer those two over the legacy status field: resource_status answers whether the preview is ready, job_status answers whether the job is still running.

Why does scene 0 matter so much?

Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. In a preview of that shape the first still fixes the avatar pose and framing, and later stills continue from it. If scene 0 is wrong, the later stills inherit the problem, so fix it before you judge anything else.

So check scene 0 for the things you cannot change after approval: avatar, background, aspect ratio and product placement. These come from avatar_handle, scene and aspect_ratio, which are structural fields.

Which review steps catch real problems?

Work through the stills in order and keep the checks short.

  • Scene 0: is the avatar the one you intended, and is the framing right for your aspect ratio (default 9:16)?
  • Product shots: does the product_image read clearly, or is it covered by the avatar?
  • Later scenes: do they feel like the same room and pose, or did the continuation drift?
  • Captions: they are stored on create and are never burned into stills, so do not judge them here.

What do I do when a still is wrong?

If only the look is off, call POST /v1/avatar-video-previews/{id}/regenerate. It reuses the stored request and refreshes only the first-frame stills, returning the same avatar_video_preview_id with a new preview-only job. If the script, video_inputs, avatar, scene or aspect ratio is wrong, create a new preview.

When it looks right, call generate-video on the preview id. Sume reuses the preview first frame when available. An empty body keeps the quality chosen at create; an optional quality overrides only the final render tier.

curl -X POST https://api.sume.com/v1/avatar-video-previews/avp_123/generate-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-preview-avp_123-final" \
  -d '{"quality":"max"}'

Is anything in the preview reused for free?

Stills are tier-independent and always reused, so approving once and then moving between standard, plus and max does not need a new preview. Admission, pre-spend, reservation, provider submit and readback all use the effective tier you pass at generate-video.

Read the exact schemas in the live OpenAPI reference linked from the Avatar video previews page. See also regenerate stills or create a new preview.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume