Avatar video request rejected? Pre-flight these 6 inputs

Catch Sume Avatar 1.0 rejections before you submit: script vs video_inputs, the 4-60 second window, silence scenes, HTTPS media and caption style. A checklist.

5 min readSume
All posts

A Sume Avatar 1.0 talking-video request is most often rejected for six reasons you can test before submitting: both script and video_inputs are present, the estimated duration falls outside 4-60 seconds, a silence scene lacks a duration, a media URL is not public HTTPS, a Korean script is paired with a Latin caption style, or inline captions are requested on a video over 60 seconds. A short local check catches all six for free.

Everything below comes from Generate avatar video and Avatar video previews. Avatar 1.0 is English-only, so the Korean case matters only for the caption rule, which the docs state for Hangul text.

The six checks

The route is POST /v1/avatar-1.0/talking-video. Launch requests use a top-level avatar_handle, or per-scene character fields inside video_inputs, to name a ready avatar.

Avatar video pre-flight checks (docs read 2026-10-10)
CheckRule in the docsIf you skip it
Script sourceProvide one of script or video_inputs, not bothRequest rejected
Duration windowEstimated target duration 4-60 seconds inclusiveShorten, or split into several jobs
Silence scenevoice.type: "silence" needs duration; script and input_text are not permittedScene rejected
Spoken scenetype: "text" with one of script or input_text, not bothScene rejected
Media URLsPublic HTTPS URLs for product_image, scene image_urlFetch fails
Caption styleKorean script with slam, punch or tiktok-green gives 400 caption_hangul_text_latin_stylePick a Hangul style or drop captions

Duration is estimated, not measured

The window applies to Sume's estimate of the finished video, so you cannot measure the length by rendering. For a single script, count words and use the rule of thumb from the script-length post. For video_inputs, set an explicit duration on each scene whenever you care about the total, since the total of all scenes must also land in 4-60 seconds.

Silence beats count toward the window. A video whose only problem is length can often be fixed by adding a silence scene at the end instead of padding the script with filler words.

Things that are not errors

Three behaviors surprise people. A caption failure is soft: the job can still succeed with a clean video_url and captions.status=failed, and inline captions create no separate billed caption job. The current execution supports one resolved avatar for each final video, and it expects the scene backgrounds to resolve to one shared scene. Resolution is 720p at this time.

If you want to approve the composition first, create an avatar video preview. Preview stills do not depend on quality, so you can raise or lower quality at generate-video without a new preview, but changing the script, video_inputs, avatar_handle, scene or aspect_ratio does need a new one.

A checklist you can wire into CI

Encode the rules once and run them on every request your system builds.

  • Exactly one of script and video_inputs is set.
  • Every silence scene has a numeric duration and no text fields.
  • Every media field starts with https:// and is publicly fetchable.
  • The estimated total is between 4 and 60 seconds, with a margin of a few seconds either side.
  • A caption style is chosen deliberately, and captions is off for anything over 60 seconds.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume