Avatar video: avatar_handle or avatar_id per scene?
Sume avatar video launch requests use avatar_handle. Per scene, a character object takes avatar_id or avatar_handle, but only one avatar is allowed per video.
Use avatar_handle at the top level of POST /v1/avatar-1.0/talking-video. Inside video_inputs a scene can name its avatar with avatar_handle or with a character object that takes avatar_id or avatar_handle. Either way, current execution supports one resolved avatar per final video, and distinct scene avatars are rejected.
Where can I put the avatar reference?
The OpenAPI schema lists three places. A top-level avatar_handle, which is optional when every scene gives its own reference. A per-scene avatar_handle. And a per-scene character object with type: "avatar" plus exactly one of avatar_id or avatar_handle.
The avatar create docs say to use the returned handle or resource id on avatar videos, and the creation response describes avatar_id as read-only, with avatar_handle used in launch generation requests.
| Where | Field | Notes |
|---|---|---|
| Top level | avatar_handle | Default for all scenes |
| Scene | avatar_handle | 2 to 31 characters, normalized without @ |
Scene character | avatar_id or avatar_handle | type must be avatar; one of the two |
| Whole video | One resolved avatar | Distinct scene avatars rejected |
What does a valid handle look like?
Handles are Instagram-style: letters, digits, underscores and periods, 2 to 30 characters after an optional leading @. Periods and underscores cannot be first, last or consecutive, and hyphens are not supported. Input is normalized to lowercase without the @.
{
"video_inputs": [
{
"id": "intro",
"character": {"type": "avatar", "avatar_handle": "@Product_Host"},
"voice": {"type": "text", "script": "Welcome back."}
}
]
}What about two presenters?
Render one video per presenter and join the clips afterwards. Giving two different avatars to two scenes in one request is rejected, so plan the cut before you submit.
If the avatar is not ready when you submit, the request returns a 409 with avatar_not_ready; wait for its resource_status to read ready.
Sources
Related posts
More in Developers
- Avatar video_inputs limits: 20 scenes, 2,000 characters each
Sume's avatar video video_inputs accepts 1 to 20 scenes, each text scene up to 2,000 characters and 60 seconds, inside the 4-60 second total window.
- Avatar video mode: sync waits 30 seconds, so use async or webhook
Sume's sync and subscribe modes wait at most 30 seconds. Avatar video usually takes longer. How to read the timed-out response and what to do next.
- Bannerbear sync API 408 after 10 seconds vs Sume mode sync
Bannerbear's sync endpoint answers 408 if the render takes over 10 seconds. Sume's mode sync waits up to 30 seconds, then returns 202 to poll.
- C2PA Conformance Explorer: vet your signer before promising credential
Before you promise clients C2PA Content Credentials, look your signing tool up in the C2PA Conformance Explorer. What it lists and how IPTC used it in 2026.
Written by Sume