Avatar video_inputs limits: 20 scenes, 2,000 characters each

Sume's avatar video video_inputs accepts 1 to 20 scenes, each text scene up to 2,000 characters and 60 seconds, inside the 4-60 second total window.

5 min readSume
All posts

On POST /v1/avatar-1.0/talking-video, video_inputs takes from 1 to 20 scenes. A spoken scene's script or input_text can be up to 2,000 characters, and a scene duration can be up to 60 seconds. Those are per-field limits; the limit that applies to the whole video is the estimated total of 4-60 seconds.

What are the exact limits?

All values below come from Sume's OpenAPI schema and the avatar video docs. The docs say scripts and multi-scene plans are accepted when Sume estimates the target duration at 4-60 seconds inclusive, and to shorten longer scripts or split them across jobs.

video_inputs limits from Sume's OpenAPI schema and docs, read 2026-10-02
FieldLimitNotes
video_inputs length1 to 20 scenesExactly one of script or video_inputs
Spoken scene text1 to 2,000 charactersOne of script or input_text
Scene duration0.001 to 60 secondsOptional on text scenes, required on silence
Total video4 to 60 seconds estimatedRejected outside the window
Scene idUp to 100 charactersOptional, for your own metadata

What happens when the plan is longer than 60 seconds?

It is not accepted as one video. Split the script into several jobs and join the clips afterwards. Inline captions have their own rule: an estimated duration above 60 seconds is rejected for them.

Because a scene can be up to 60 seconds long on its own, check the sum of your scene durations rather than each scene alone.

{
  "avatar_handle": "product_host",
  "aspect_ratio": "9:16",
  "video_inputs": [
    {"id": "hook", "voice": {"type": "text", "script": "One selfie, one full video.", "duration": 3}},
    {"id": "beat", "voice": {"type": "silence", "duration": 2}},
    {"id": "cta", "voice": {"type": "text", "input_text": "Pick a template and drop in your photo.", "duration": 5}}
  ]
}

What stays true across scenes?

These rules hold for every scene in one request.

  • One resolved avatar per final video: distinct scene avatars are rejected.
  • Scene backgrounds must resolve to one shared scene.
  • Background images are public HTTPS URLs; video backgrounds are not supported.
  • The first-frame preview route uses the same fields and the same 4-60 second window.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume