Approval gate for AI avatar video: review stills before rendering
Add a sign-off step before publishing synthetic video: Sume's avatar video preview returns first-frame stills, and a later tier change needs no new preview.
Short answer
You can make a review gate out of Sume's avatar video preview. It creates first-frame stills without starting the full talking-video render, so a reviewer can approve the look, the presenter and the scene before anyone pays for the video or publishes it. Because platforms such as YouTube ask creators to judge whether realistic synthetic content needs a label, having a named person look at the stills first is a simple control.
The four calls
The preview resource has four routes. Together they make a loop of create, review, optionally regenerate, and render.
| Step | Route | Purpose |
|---|---|---|
| Create | POST /v1/avatar-video-previews | First-frame stills, no full render |
| Read | GET /v1/avatar-video-previews/:id | preview_image_url and scene_previews |
| Redo | POST /v1/avatar-video-previews/:id/regenerate | New stills from the stored request |
| Render | POST /v1/avatar-video-previews/:id/generate-video | Start the real video |
What the reviewer sees
When the preview is ready, the public-safe fields include preview_image_url, the primary still, and scene_previews, one still per input scene. For multi-scene previews with a shared scene, later stills are pose-anchored continuations of the first frame. The resource reports resource_status and job_status; the docs prefer these to the older status field. The reviewer checks the face, the setting, and whether the result could be mistaken for real footage of a real person.
- Is the presenter a real, identifiable person?
- Is the scene a real place that has been altered?
- Does the frame look like something that actually happened?
- Where will the disclosure appear?
Changing your mind after approval
Preview stills are tier-independent, and Sume reuses them. At generate-video, an empty body keeps the quality you chose at create, and an optional quality overrides only the final render tier. If you approve and then downgrade or upgrade, you do not need a new preview. Changes to structural fields such as script, video_inputs, avatar_handle, scene or aspect_ratio do need a new preview, which is the right behavior for a gate: a different video needs a new approval.
| Change | New preview needed |
|---|---|
| quality at generate-video | No |
| script or video_inputs | Yes |
| avatar_handle | Yes |
| scene | Yes |
| aspect_ratio | Yes |
Record the decision
Keep the avatar_video_preview_id in your own log with the reviewer's name, the date and the label decision. Captions you store at preview create apply when you render, so the text a reviewer saw is the text that ships. Sume does not decide the label for you; the YouTube help page leaves that judgment to the creator, and your log shows how you made it.
A small team can run this with no new tooling: the person who writes the script creates the preview, a second person reads the stills and the script together, and only then does the first person call generate-video. If the reviewer rejects the look, regenerate the stills from the stored request instead of rebuilding it, and review again. The loop costs a preview, not a full video, until the approval is given.
Sources
Related posts
More in Sume Avatar 1.0
- Griffin-Lite 26 of 54: error bars and a viewer test for avatars
Tavus says 48% of 54 people took Griffin-Lite for real. With 54 viewers the 95% range is 35% to 61%. A Python check and a test plan for Sume avatar clips.
- Can you buy Tavus Griffin-Lite? Presenter video options today
Tavus Griffin-Lite is a research preview for trusted testers, not a product customers can buy. What Sume ships for presenter videos in the meantime.
- Vidu Q4 audio references vs a scripted avatar: which keeps a voice?
Vidu Q4 Preview takes up to 3 audio references; Sume does not list Vidu. How a scripted avatar video keeps a presenter's face and words consistent.
- What Sume Avatar 1.0 does not do: eight limits to check first
No streaming, no interruption, English-only speech in code, 720p, 4 to 60 seconds, one avatar per video. The limits of Sume Avatar 1.0 in one table.
Written by Sume