Avatar video script over 60 seconds: what Sume does and how to split
Sume avatar videos accept 4 to 60 seconds, and inline captions stop at 60. How to split a longer script, add a silent beat, and price the parts.
Sume will not render it as one video. A script or multi-scene plan is accepted only when Sume estimates the video at 4 to 60 seconds inclusive, and the docs tell you to shorten the script or split it into multiple jobs. Inline captions have their own cap: an estimated duration above 60 seconds is rejected for them.
How to split a 3-minute script
A 3-minute script at about 150 spoken words a minute is roughly 450 words. Four parts of about 112 words each land near 45 seconds.
- Cut at topic changes, not at word counts. Aim for parts of 45 to 55 seconds so a slow read still fits.
- Give every part the same avatar_handle, aspect_ratio and quality so the parts match.
- Use a different Idempotency-Key per part.
- Join the parts afterward in a timeline render.
Use scenes for pacing inside 60 seconds
With video_inputs you can order scenes for hook, demo and close. A scene with voice.type silence is a non-speaking beat: duration is required, script is not allowed. Total planned duration must still fit the 4 to 60 second window. Current execution supports one avatar per final video and scene backgrounds that resolve to one shared scene.
What four parts cost
Four 45-second parts on plus without a product image: 4 x 45 x $0.245 = $44.10. On standard it is 4 x 45 x $0.184 = $33.12. Add $0.10 per output minute for a Timeline 1.0 render to join them, rounded up to the minute: three minutes adds $0.30.
| Tier | Per part | Four parts | With timeline join |
|---|---|---|---|
| standard | $8.28 | $33.12 | $33.42 |
| plus | $11.03 | $44.10 | $44.40 |
A silent beat in the request
{
"avatar_handle": "your_avatar",
"aspect_ratio": "9:16",
"video_inputs": [
{ "id": "hook", "voice": { "type": "text", "script": "Here is the problem in one sentence." }, "background": { "type": "prompt", "prompt": "Bright home office" } },
{ "id": "beat", "voice": { "type": "silence", "duration": 2 }, "background": { "type": "prompt", "prompt": "Bright home office" } },
{ "id": "fix", "voice": { "type": "text", "script": "And here is the fix." }, "background": { "type": "prompt", "prompt": "Bright home office" } }
]
}Sources
Related posts
More in Sume Avatar 1.0
- Azure photo avatar is 512x512; Sume avatar video is 720p
Azure's photo avatar renders head-only at 512x512 and 25 fps. Sume's avatar video offers 720p in five aspect ratios, from a prompt, traits or a photo.
- Azure avatar batch: 20-minute videos, 200 jobs, vs Sume
Azure's text to speech avatar batch API allows 20-minute outputs and 200 concurrent jobs per Speech resource. Sume avatar videos run 4-60 seconds per job.
- How to check a face swap result: audio, length and stills
Check a Sume face swap before publishing: confirm the audio stream, compare length with the source, and sample stills with video inspect. Free probe.
- Create an AI spokesperson avatar from a prompt or photo
Sume Avatar 1.0 creates a reusable avatar at a flat $0.95 from a text prompt, a photo, or props. How to make one before you generate talking videos.
Written by Sume