Avatar preview final render has no webhook_url: poll the job
The preview generate-video call takes only quality, so it has no webhook_url. Poll the job, or use the direct avatar endpoint for a webhook-driven render.
Starting the final render from an approved preview does not accept a webhook_url; the generate-video body takes quality and nothing else, so you follow that job by polling its status URL. If you need a signed webhook for the render, create the avatar video directly instead of going through the preview stage.
This is from Avatar video previews and the API schema in this repository, read 2026-10-10, plus Webhooks.
Which path takes which fields
Both paths produce an avatar video; they differ in what you can attach.
| Step | Endpoint | What it accepts | Follow it by |
|---|---|---|---|
| Preview create | POST /v1/avatar-video-previews | Same fields as Avatar Video: script or video_inputs, scene, quality, aspect ratio, title, captions | Poll the preview job, then read the preview resource |
| Regenerate stills | POST /v1/avatar-video-previews/:id/regenerate | Reuses the stored request | Poll the job |
| Final render | POST /v1/avatar-video-previews/:id/generate-video | quality only | Poll the job |
| Direct render | POST /v1/avatar-1.0/talking-video | Full body, async, sync, subscribe or webhook mode | Poll or receive a signed event |
Polling the final render
The response carries the job's status and result URLs, as for any generation. Poll the status URL until terminal is true, then read the result when result_ready is true. On the preview resource, prefer resource_status and job_status over the legacy status field.
A reasonable cadence is a few seconds early on and slower after, with a hard cap on total wait so a stuck poll cannot run forever. Avatar videos are minutes-long jobs at the top of the range, not seconds.
When to pick which
- A person approves the first frame before spending: use previews and poll the render.
- A pipeline renders hundreds with no approval: use the direct endpoint with a webhook and verify the signature.
- You want both: use the preview stage for the first of a series, then render the rest directly with the same avatar and settings.
Quality is chosen at the end
Because preview stills are tier-independent, the quality choice happens at generate-video. Default quality on preview create is plus, so a script left at defaults is not priced the same as the standard tier. Set it deliberately on the final call.
A polling loop that behaves
Poll GET /v1/jobs/{id}/status with the same bearer key you used to start the job. Read terminal; when it is true and result_ready is true, fetch the result. If terminal is true and result_ready is false, the job failed or was canceled, and the result route will answer 409 job_not_completed, so read the failure from the job record.
Wait a few seconds between polls, cap the total wait, and surface a clear message when you hit the cap. The job itself keeps running after your loop gives up, so a later check by id still finds it.
Why this matters for a pipeline
A pipeline built entirely around webhooks will silently miss the preview route. If your design says every job reports to my endpoint, add one exception: preview-originated renders are polled. A small table of job ids to wait on, checked by a timer, covers it.
Keep the preview resource id next to the job id in that table. The job tells you whether the render finished; the preview resource tells you which approved frame and request it came from, which is what a reviewer will ask about later.
Edge cases
If you regenerate stills before starting a render, nothing is billed for the video yet. Once generate-video is called, treat the render as committed: a job that has started cannot be canceled, and the API answers with a conflict if you try. Poll to the end, read the result, and only then decide whether to approve it.
Sources
Related posts
More in Sume Avatar 1.0
- Face swap or talking video: which Avatar endpoint fits?
Pick Sume Avatar face swap (Beta) or Avatar talking video by what you have: footage with audio, or a script. Inputs, limits and polling side by side.
- Gradium's 329 character voices vs voices of Sume avatars
Gradium lists 329 character voices. On Sume, a spoken character is an avatar with a ready voice; Avatar 1.0 is English-only. What that means for casting.
- Moving Company Spokesperson Ad: A 20-Second Avatar for About $4.90
A 20-second plus-quality Avatar 1.0 spokesperson clip costs about $4.90 on Sume, at 98 cents per 4 seconds. English only, with captions at $0.20 extra.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume