Can a Sume avatar speak my own recording? Voice types explained

Avatar 1.0 talking video takes a script or per-scene voice blocks of type text or silence. What the docs list, what they do not, and how to plan a pause.

3 min readSume
All posts

The Avatar 1.0 talking-video docs describe speech as written text: a top-level script, or ordered video_inputs where each scene's voice block is either type: "text" or type: "silence". They do not document an audio-file input for this route, so plan on text in and a spoken avatar out (Generate avatar video).

If your workflow depends on a specific recording, check the live OpenAPI before you build around it, and do not assume the route accepts one.

What each voice type means

Per-scene voice rules in the docs
Voice typeRequiredNot allowed
textscript or input_text (one of them), durationboth script and input_text
silencedurationscript, input_text

A scene list with a pause

The pause is a real scene with a duration, so it counts toward the 4 to 60 second window.

{
  "avatar_handle": "product_host",
  "quality": "plus",
  "video_inputs": [
    { "id": "hook", "voice": { "type": "text", "script": "Wait, one selfie became a whole video?", "duration": 3 } },
    { "id": "beat", "voice": { "type": "silence", "duration": 2 } },
    { "id": "cta",  "voice": { "type": "text", "input_text": "Pick a template and drop in your photo.", "duration": 4 } }
  ]
}

Cost of that plan

Nine seconds in total, billed at the tier rate with no product image.

Sume list price, read 2026-10-08
Tier9 s
standard$1.656
plus$2.205
max$4.95

Other constraints

The current execution supports one resolved avatar for each final video, and the scene backgrounds are expected to resolve to one shared scene. Total planned duration must stay in the 4 to 60 second window.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume