A 170-word avatar script in 34-word sentences plans 70 s: refused
Why a 170-word script of five 34-word sentences needs 70 seconds in Avatar 1.0, past the 60-second cap, and how to split it into two 35-second videos.
A script of 170 words is not safe under the 60-second limit, because the limit is on planned seconds, not on words. If the script is five sentences of 34 words, each sentence is split into two 17-word parts of 7 seconds, which makes ten clips and 70 seconds in total. The planner then refuses it with the message "Avatar video script is estimated at 70 seconds; maximum is 60 seconds. Shorten the script or split it into multiple videos." Even a plain division gives 170 / 2.8 = 60.7 seconds, which is already over; the split rounding adds the rest.
Where the extra ten seconds come from
In the avatar-workflows package of the Sume repo, speech is estimated at 2.8 words per second, a clip is 4 to 12 seconds, and a sentence of more than 33 words is cut in equal parts. A 34-word sentence becomes 17 + 17. Each part is rounded up to a whole second: ceil(17 / 2.8) = 7. Two parts make 14 seconds, where 34 words would need 12.1. Five such sentences make 70.
The Avatar videos docs describe the outcome: the talking video accepts 4 to 60 seconds. They do not publish the planner constants, so treat the figures here as the behavior I measured on 2026-10-08, not as a promise.
| Script shape | Clips | Planned seconds | Result |
|---|---|---|---|
| 5 sentences x 34 words | 10 | 70 | Refused, over 60 |
| 10 sentences x 16 words, plus 10 words | 6 | 64 | Refused, over 60 |
| Two videos of 5 sentences x 17 words | 5 each | 35 each | Accepted, two jobs |
Why shorter sentences do not rescue it
It is tempting to rewrite the long sentences and keep one video. Ten sentences of 16 words pack two to a clip, 32 words each, planned at 12 seconds. That is five clips and 60 seconds for 160 words, and the last 10 words add a sixth clip of 4 seconds, for 64. Any script above about 168 words cannot fit, because 60 seconds times 2.8 words per second is 168, and the planner rounds each clip up. So the fix is a second video.
Fix: split into two videos
Split the script at a section break into two halves of about 85 words, so each plans at 35 seconds when the sentences are 17 words.
- Count words per sentence and flag any over 33.
- Estimate total seconds as the sum of ceil(words / 2.8) per clip, not total words / 2.8.
- Keep the total at or under 60 for each job.
- Split at a topic change, not mid-argument, and use the same avatar handle in both jobs.
- Join the two outputs with
POST /v1/timeline-1.0/render, which bills $0.10 per output minute, rounded up.
What Sume does not do
Sume does not trim your script or lengthen the window to fit. The refusal comes before any render, so a script that is too long costs nothing, but it also gives you no partial result. If you send scenes instead of a script, the same cap applies to the sum of scene durations, with a message that talks about scenes rather than a script.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video with a product image: the per-second price gap
On Sume Avatar 1.0 a product_image adds $0.010 to $0.030 per second by tier. For a 30-second spokesperson clip that is $0.30 to $0.90. Arithmetic inside.
- Avatar video succeeded but captions failed: what to do on Sume
On Sume a caption-stage failure is soft: the avatar job can still succeed with a clean video_url and captions.status=failed. How to check it and add captions.
- Avatar video with a product image: +$0.01 to +$0.03 a second
Adding a product image to a Sume Avatar 1.0 video raises the per-second rate by $0.01 on standard, $0.013 on plus and $0.03 on max. Totals at 15, 30 and 60 s.
- Sume TTS voice then H3 Max lip sync: 40 seconds in 3 slices
H3 Max lip-sync accepts 5 to 14.8 s of audio. Split a 40 second TTS script into three sentence-based slices at $4.00 on 768p, and what to check.
Written by Sume