Avatar video_inputs voice rules: text, silence, script vs input_text
Each video_inputs scene takes a voice of type text or silence. Text needs exactly one of script or input_text; silence needs duration and forbids both.
In an Avatar Video video_inputs scene, voice.type is either "text" or "silence". A text scene needs exactly one of script or input_text; a silence scene needs duration and must not carry script or input_text at all.
These rules come from the Generate avatar video guide, read 2026-10-02. They apply to POST /v1/avatar-1.0/talking-video and to the preview route that takes the same body.
What are the two voice types?
Use video_inputs instead of a top-level script when you want hooks, demos or quiet beats in one composed video. You provide exactly one of script or video_inputs on the request; sending both is a different mistake from the one this post covers.
| voice.type | Required | Not allowed |
|---|---|---|
| text | Exactly one of script or input_text | Sending both, or neither |
| silence | duration | script and input_text |
Do I use script or input_text?
The guide treats them as two names for the spoken words of one scene and requires exactly one. Its own multi-scene example uses script in the hook and input_text in the call to action, so both spellings are valid; pick one per scene and stay consistent in your code so a validator is simple.
A text scene can also carry a duration. The guide's example sets 3 seconds for the hook and 5 for the call to action. What matters at the request level is that Sume estimates the whole video at 4 to 60 seconds inclusive.
What does a silence beat look like?
A silence scene is a non-speaking beat, useful for letting a product demo breathe. Here is the shape, trimmed from the guide's example:
{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "Wait, one selfie became a video?", "duration": 3}},
{"id": "demo", "voice": {"type": "silence", "duration": 4}},
{"id": "cta", "voice": {"type": "text", "input_text": "Pick a template, drop in your photo.", "duration": 5}}
]
}The guide's full example also gives every scene a background prompt; this trimmed version leaves it out for space. See the background post for that field.
Which limits still apply to a multi-scene plan?
Total planned duration must still land in the 4 to 60 second window, silence included. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so video_inputs is not a way to cast two hosts in one clip.
For per-scene character fields, see avatar id or handle per scene.
How do I catch the mistake before paying?
Create a preview first: it takes the same fields, produces one still per scene, and does not start the full render. A malformed voice block should fail at that step. Structural fields such as video_inputs cannot change at generate-video, so a script fix means a new preview.
Sume does not guess which of your two fields you meant. Build the scene list from one function that emits one key, and the rule never bites.
Should I validate on my side too?
If your app builds scene lists from user input, enforce the rule on your side so users see a readable message before Sume sees the request. The check is short per scene: text scenes have exactly one of the two keys, silence scenes have a duration and neither key. Do not invent extra rules; keep to what the guide states.
Keep scene id values unique and descriptive, such as hook, demo and cta. The guide's example uses them, and they make a preview's scene_previews[] stills easy to match to the scene that produced them.
How should I time a silence beat?
The guide only requires a duration; it does not state a minimum or maximum for a single beat. What it does fix is the whole video: 4 to 60 seconds inclusive. If you plan a 4 second silence between a 3 second hook and a 5 second closer, that is 12 seconds in total, comfortably inside the window.
A silence beat is not a pause inside a sentence. Because it is its own scene, use it between thoughts, not mid-sentence. For rhythm inside a spoken line, write shorter sentences in the script instead.
Remember that a scene with a product on screen and no speech is a good use of silence: the viewer reads the product while the avatar stays quiet.
Sources
Related posts
More in Sume Avatar 1.0
- Sume avatar video stuck? Read the job events before retrying
Use GET /v1/jobs/{id}/events to see whether an avatar video is queued, started, submitted or failed before you cancel, wait, or resubmit and pay twice.
- Avatar video soundtrack: send prompt or audio_url, volume 0.05-0.4
Sume's Avatar Video package accepts a soundtrack with exactly one of prompt or audio_url and a volume from 0.05 to 0.4, default 0.15. What each costs and does.
- Avatar video preview: approve the first frame, then pick quality
Sume avatar video previews are tier-independent: approve stills, then render at standard, plus or max. What can change at generate-video, and what cannot.
- Avatar preview resource_status vs job_status: which field to poll
Avatar video previews return resource_status and job_status next to a legacy status field. Read resource_status for readiness and job_status for polling.
Written by Sume