Write an avatar script that sounds authentic in 4 to 60 seconds

HeyGen says 64.6% trust avatars that sound authentic. Draft a Sume avatar script in the 4-60 second window with scene beats and a silence gap.

5 min readSume
All posts

An avatar script that sounds authentic is short, specific and spoken aloud before you submit it: one idea per scene, plain sentences, and a total that Sume estimates at 4 to 60 seconds. If it does not, shorten it or split it into several jobs, because longer scripts are not accepted.

The reason to care is a line in HeyGen's The State of AI Avatars 2026: 64.6% of respondents say they trust avatars that sound authentic, and 57.1% value personal, high-quality content. It is a vendor survey, but the writing advice below holds whatever the number is, and every field comes from Sume's avatar video guide.

How long can the script be?

Provide exactly one of script or video_inputs. Sume estimates the target duration, and the accepted window is 4-60 seconds inclusive. With video_inputs, the total planned duration has to land in the same window. There is no way to extend it in a single request, so a two-minute message becomes two or more jobs.

When should you use scenes instead of one script?

Use ordered video_inputs when the clip has beats: a hook, a demo, a call to action. Each scene has a voice that is either spoken text (exactly one of script or input_text, plus an optional duration) or silence, which needs a duration and no text. A silence beat is how you give the viewer time to look at a product instead of talking over it.

Scene voice types in Avatar video_inputs (read 2026-10-03)
Voice typeRequired fieldsNot allowed
textexactly one of script or input_textboth at once
silencedurationscript and input_text

What does an authentic three-beat script look like?

Write like you talk to one customer: a question, one concrete thing, one next step. The example below uses one shared background prompt for all scenes and keeps the demo beat silent. Note that current execution supports one resolved avatar per final video and expects the scene backgrounds to resolve to one shared scene.

{
  "avatar_handle": "shop_owner",
  "aspect_ratio": "9:16",
  "video_inputs": [
    {"id": "hook",
     "voice": {"type": "text", "script": "Ever lose a Saturday to packing orders?", "duration": 3},
     "background": {"type": "prompt", "prompt": "Small workshop, natural light"}},
    {"id": "demo",
     "voice": {"type": "silence", "duration": 4},
     "background": {"type": "prompt", "prompt": "Small workshop, natural light"}},
    {"id": "cta",
     "voice": {"type": "text", "input_text": "Order by Thursday and it ships Friday.", "duration": 4},
     "background": {"type": "prompt", "prompt": "Small workshop, natural light"}}
  ]
}

What makes a script sound read rather than spoken?

Read it aloud once. Cut anything you would not say to a friend, such as stacked adjectives and slogans. Replace a list of five claims with the one you can prove. Put the number or the name in the sentence instead of a vague promise.

Keep scene text short enough for its duration: a long sentence squeezed into three seconds is the quickest way to make speech feel rushed. If the estimate lands outside 4-60 seconds the request is refused, so trim before you pay for a render.

What does Sume not do here?

Sume does not rewrite your script and does not judge whether it is authentic. It also generates a presenter rather than recording you, so the authenticity in the survey comes from the words and the framing you choose, not from the model. Use a first-frame preview to check framing before the full render.

How do you check the script before paying for a render?

Create an avatar video preview with the same script or video_inputs. It returns first-frame stills without starting the full render, and the final-render quality can still be chosen at generate-video. If you then change the text, the preview is no longer valid for it: a changed script or video_inputs needs a new preview.

Inline captions are optional and are applied after generation using the spoken text, so they follow your script automatically. Estimated duration above 60 seconds is rejected for inline captions too.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume