Avatar video request rejected? Pre-flight these 6 inputs
Catch Sume Avatar 1.0 rejections before you submit: script vs video_inputs, the 4-60 second window, silence scenes, HTTPS media and caption style. A checklist.
A Sume Avatar 1.0 talking-video request is most often rejected for six reasons you can test before submitting: both script and video_inputs are present, the estimated duration falls outside 4-60 seconds, a silence scene lacks a duration, a media URL is not public HTTPS, a Korean script is paired with a Latin caption style, or inline captions are requested on a video over 60 seconds. A short local check catches all six for free.
Everything below comes from Generate avatar video and Avatar video previews. Avatar 1.0 is English-only, so the Korean case matters only for the caption rule, which the docs state for Hangul text.
The six checks
The route is POST /v1/avatar-1.0/talking-video. Launch requests use a top-level avatar_handle, or per-scene character fields inside video_inputs, to name a ready avatar.
| Check | Rule in the docs | If you skip it |
|---|---|---|
| Script source | Provide one of script or video_inputs, not both | Request rejected |
| Duration window | Estimated target duration 4-60 seconds inclusive | Shorten, or split into several jobs |
| Silence scene | voice.type: "silence" needs duration; script and input_text are not permitted | Scene rejected |
| Spoken scene | type: "text" with one of script or input_text, not both | Scene rejected |
| Media URLs | Public HTTPS URLs for product_image, scene image_url | Fetch fails |
| Caption style | Korean script with slam, punch or tiktok-green gives 400 caption_hangul_text_latin_style | Pick a Hangul style or drop captions |
Duration is estimated, not measured
The window applies to Sume's estimate of the finished video, so you cannot measure the length by rendering. For a single script, count words and use the rule of thumb from the script-length post. For video_inputs, set an explicit duration on each scene whenever you care about the total, since the total of all scenes must also land in 4-60 seconds.
Silence beats count toward the window. A video whose only problem is length can often be fixed by adding a silence scene at the end instead of padding the script with filler words.
Things that are not errors
Three behaviors surprise people. A caption failure is soft: the job can still succeed with a clean video_url and captions.status=failed, and inline captions create no separate billed caption job. The current execution supports one resolved avatar for each final video, and it expects the scene backgrounds to resolve to one shared scene. Resolution is 720p at this time.
If you want to approve the composition first, create an avatar video preview. Preview stills do not depend on quality, so you can raise or lower quality at generate-video without a new preview, but changing the script, video_inputs, avatar_handle, scene or aspect_ratio does need a new one.
A checklist you can wire into CI
Encode the rules once and run them on every request your system builds.
- Exactly one of
scriptandvideo_inputsis set. - Every silence scene has a numeric
durationand no text fields. - Every media field starts with
https://and is publicly fetchable. - The estimated total is between 4 and 60 seconds, with a margin of a few seconds either side.
- A caption style is chosen deliberately, and
captionsis off for anything over 60 seconds.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar preview final render has no webhook_url: poll the job
The preview generate-video call takes only quality, so it has no webhook_url. Poll the job, or use the direct avatar endpoint for a webhook-driven render.
- Face swap or talking video: which Avatar endpoint fits?
Pick Sume Avatar face swap (Beta) or Avatar talking video by what you have: footage with audio, or a script. Inputs, limits and polling side by side.
- Gradium's 329 character voices vs voices of Sume avatars
Gradium lists 329 character voices. On Sume, a spoken character is an avatar with a ready voice; Avatar 1.0 is English-only. What that means for casting.
- Moving Company Spokesperson Ad: A 20-Second Avatar for About $4.90
A 20-second plus-quality Avatar 1.0 spokesperson clip costs about $4.90 on Sume, at 98 cents per 4 seconds. English only, with captions at $0.20 extra.
Written by Sume