Avatar video_inputs limits: 20 scenes, 2,000 characters each
Sume's avatar video video_inputs accepts 1 to 20 scenes, each text scene up to 2,000 characters and 60 seconds, inside the 4-60 second total window.
On POST /v1/avatar-1.0/talking-video, video_inputs takes from 1 to 20 scenes. A spoken scene's script or input_text can be up to 2,000 characters, and a scene duration can be up to 60 seconds. Those are per-field limits; the limit that applies to the whole video is the estimated total of 4-60 seconds.
What are the exact limits?
All values below come from Sume's OpenAPI schema and the avatar video docs. The docs say scripts and multi-scene plans are accepted when Sume estimates the target duration at 4-60 seconds inclusive, and to shorten longer scripts or split them across jobs.
| Field | Limit | Notes |
|---|---|---|
video_inputs length | 1 to 20 scenes | Exactly one of script or video_inputs |
| Spoken scene text | 1 to 2,000 characters | One of script or input_text |
Scene duration | 0.001 to 60 seconds | Optional on text scenes, required on silence |
| Total video | 4 to 60 seconds estimated | Rejected outside the window |
Scene id | Up to 100 characters | Optional, for your own metadata |
What happens when the plan is longer than 60 seconds?
It is not accepted as one video. Split the script into several jobs and join the clips afterwards. Inline captions have their own rule: an estimated duration above 60 seconds is rejected for them.
Because a scene can be up to 60 seconds long on its own, check the sum of your scene durations rather than each scene alone.
{
"avatar_handle": "product_host",
"aspect_ratio": "9:16",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "One selfie, one full video.", "duration": 3}},
{"id": "beat", "voice": {"type": "silence", "duration": 2}},
{"id": "cta", "voice": {"type": "text", "input_text": "Pick a template and drop in your photo.", "duration": 5}}
]
}What stays true across scenes?
These rules hold for every scene in one request.
- One resolved avatar per final video: distinct scene avatars are rejected.
- Scene backgrounds must resolve to one shared scene.
- Background images are public HTTPS URLs; video backgrounds are not supported.
- The first-frame preview route uses the same fields and the same 4-60 second window.
Sources
Related posts
More in Developers
- Avatar video mode: sync waits 30 seconds, so use async or webhook
Sume's sync and subscribe modes wait at most 30 seconds. Avatar video usually takes longer. How to read the timed-out response and what to do next.
- Bannerbear sync API 408 after 10 seconds vs Sume mode sync
Bannerbear's sync endpoint answers 408 if the render takes over 10 seconds. Sume's mode sync waits up to 30 seconds, then returns 202 to poll.
- C2PA Conformance Explorer: vet your signer before promising credential
Before you promise clients C2PA Content Credentials, look your signing tool up in the C2PA Conformance Explorer. What it lists and how IPTC used it in 2026.
- C2PA Interim Trust List frozen: why Content Credentials look untrusted
The C2PA Interim Trust List was frozen on 1 January 2026. What that means when a validator flags credentials, and what to check on files you deliver.
Written by Sume