Avatar video preview: approve the first frame, then pick quality

Sume avatar video previews are tier-independent: approve stills, then render at standard, plus or max. What can change at generate-video, and what cannot.

5 min readSume
All posts

With Sume's avatar video previews, you approve first-frame stills once and then choose the quality tier of the final render when you call generate-video. Preview stills are tier-independent and are always reused, so you can approve on a cheap check and then render at max, or the other way round, without making a new preview. What you cannot change at that step is anything structural: script, video_inputs, avatar_handle, scene and aspect_ratio need a new preview.

That split is useful when a stakeholder has to sign off on framing, but nobody wants to pay for the best render before they do.

The four calls

The preview resource has its own endpoints.

From Avatar video previews, read 2026-10-01.
CallWhat it does
POST /v1/avatar-video-previewsCreates the first-frame stills job; same body as a talking video, quality defaults to plus
GET /v1/avatar-video-previews/:idReads preview_image_url and scene_previews[] when ready
POST /v1/avatar-video-previews/:id/regenerateRefreshes only the stills from the stored request
POST /v1/avatar-video-previews/:id/generate-videoStarts the full render; optional quality overrides the final tier only

A worked flow

Create the preview with quality: "standard", review the still, and if the composition is right call generate-video with an empty body to keep the same tier, or with {"quality": "max"} to upgrade. Admission, the pre-spend check, the ledger reservation, the provider submit and the readback all use the effective tier, so the price you are quoted is the price of the override.

If the still is wrong, call regenerate. It reuses the stored request, returns the same avatar_video_preview_id and runs a new preview-only job.

curl -X POST https://api.sume.com/v1/avatar-video-previews/avp_123/generate-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avp-123-final" \
  -d '{"quality": "max"}'

Cost of the choice

For a 20-second clip with no product image, using the current rate card (confirm live with GET /v1/catalog):

Arithmetic from the Sume rate card for talking-video, no product image; 20 seconds is rate times 20.
Final tierPer second20 seconds
standard$0.184$3.68
plus$0.245$4.90
max$0.55$11.00

Other things to know

  • Captions stored on preview create are applied only at generate-video; preview stills are never captioned.
  • For multi-scene plans, later scene stills are pose-anchored continuations of the first frame.
  • Prefer resource_status and job_status over the legacy status field when deciding what is ready.
  • The duration window is the same as a normal avatar video: an estimated 4 to 60 seconds.

Limits

A preview shows a first frame, not motion or lip-sync, so it will not tell you whether a long word is awkward on camera. Treat it as a framing and look check, then read the first render before you scale to many videos.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume