Avatar video_inputs voice rules: text, silence, script vs input_text

Each video_inputs scene takes a voice of type text or silence. Text needs exactly one of script or input_text; silence needs duration and forbids both.

5 min readSume
All posts

In an Avatar Video video_inputs scene, voice.type is either "text" or "silence". A text scene needs exactly one of script or input_text; a silence scene needs duration and must not carry script or input_text at all.

These rules come from the Generate avatar video guide, read 2026-10-02. They apply to POST /v1/avatar-1.0/talking-video and to the preview route that takes the same body.

What are the two voice types?

Use video_inputs instead of a top-level script when you want hooks, demos or quiet beats in one composed video. You provide exactly one of script or video_inputs on the request; sending both is a different mistake from the one this post covers.

voice.type rules per scene (read 2026-10-02)
voice.typeRequiredNot allowed
textExactly one of script or input_textSending both, or neither
silencedurationscript and input_text

Do I use script or input_text?

The guide treats them as two names for the spoken words of one scene and requires exactly one. Its own multi-scene example uses script in the hook and input_text in the call to action, so both spellings are valid; pick one per scene and stay consistent in your code so a validator is simple.

A text scene can also carry a duration. The guide's example sets 3 seconds for the hook and 5 for the call to action. What matters at the request level is that Sume estimates the whole video at 4 to 60 seconds inclusive.

What does a silence beat look like?

A silence scene is a non-speaking beat, useful for letting a product demo breathe. Here is the shape, trimmed from the guide's example:

{
  "avatar_handle": "sume_clawra",
  "aspect_ratio": "9:16",
  "quality": "plus",
  "video_inputs": [
    {"id": "hook", "voice": {"type": "text", "script": "Wait, one selfie became a video?", "duration": 3}},
    {"id": "demo", "voice": {"type": "silence", "duration": 4}},
    {"id": "cta", "voice": {"type": "text", "input_text": "Pick a template, drop in your photo.", "duration": 5}}
  ]
}

The guide's full example also gives every scene a background prompt; this trimmed version leaves it out for space. See the background post for that field.

Which limits still apply to a multi-scene plan?

Total planned duration must still land in the 4 to 60 second window, silence included. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so video_inputs is not a way to cast two hosts in one clip.

For per-scene character fields, see avatar id or handle per scene.

How do I catch the mistake before paying?

Create a preview first: it takes the same fields, produces one still per scene, and does not start the full render. A malformed voice block should fail at that step. Structural fields such as video_inputs cannot change at generate-video, so a script fix means a new preview.

Sume does not guess which of your two fields you meant. Build the scene list from one function that emits one key, and the rule never bites.

Should I validate on my side too?

If your app builds scene lists from user input, enforce the rule on your side so users see a readable message before Sume sees the request. The check is short per scene: text scenes have exactly one of the two keys, silence scenes have a duration and neither key. Do not invent extra rules; keep to what the guide states.

Keep scene id values unique and descriptive, such as hook, demo and cta. The guide's example uses them, and they make a preview's scene_previews[] stills easy to match to the scene that produced them.

How should I time a silence beat?

The guide only requires a duration; it does not state a minimum or maximum for a single beat. What it does fix is the whole video: 4 to 60 seconds inclusive. If you plan a 4 second silence between a 3 second hook and a 5 second closer, that is 12 seconds in total, comfortably inside the window.

A silence beat is not a pause inside a sentence. Because it is its own scene, use it between thoughts, not mid-sentence. For rhythm inside a spoken line, write shorter sentences in the script instead.

Remember that a scene with a product on screen and no speech is a good use of silence: the viewer reads the product while the avatar stays quiet.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume