Five 20-word avatar sentences make five 8-second clips, not three

Why five 20-word sentences plan as five 8-second clips (40 s) in Avatar 1.0, and how pairing shorter sentences saves clips.

4 min readSume
All posts

Five sentences of 20 words each, 100 words in all, plan as five clips of 8 seconds each, a total of 40 seconds, because two of them together would be 40 words and the clip ceiling is 33. Each sentence is therefore closed on its own. If you shorten the sentences to 16 words, two fit in one clip, and the same ideas take fewer clips.

The packing rule applied

The numbers on this page come from running the avatar-workflows package in the Sume repo on constructed scripts, and from the Avatar videos docs for the 4 to 60 second window. The chunker is an internal planner: it decides how many clips your script becomes, and the public docs describe the outcome rather than the algorithm, so treat the exact constants as current behavior that can change.

The planner never splits a sentence that fits, and never overfills a clip. With 20-word sentences the first sentence is 20 words, the second would bring it to 40, which is over 33, so the clip is closed. This repeats five times. Each clip then lasts ceil(20 / 2.8) = 8 seconds.

Same ideas, different sentence length (read 2026-10-08)
ScriptWordsClipsPlanned seconds
5 sentences x 20 words100540
6 sentences x 16 words (2 per clip)96312 each, 36 total
10 sentences x 10 words (3 per clip)100411, 11, 11 and 4, 37 total

A check you can run

The mirror below plans a script that you paste in. The script prints the words and planned seconds for each clip; compare the output to your budget.

import math, re
def chunks(script):
    maxw, out, cur = 33, [], []
    for s in re.findall(r"[^.!?]+[.!?]?", script):
        w = s.split()
        if len(w) > maxw:
            if cur:
                out.append(cur)
                cur = []
            n = math.ceil(len(w) / maxw)
            k, r = divmod(len(w), n)
            i = 0
            for j in range(n):
                size = k + (j < r)
                out.append(w[i:i + size])
                i += size
        elif cur and len(cur) + len(w) > maxw and len(cur) >= 7:
            out.append(cur)
            cur = w
        else:
            cur = cur + w
    if cur:
        out.append(cur)
    return [(len(c), min(12, max(4, math.ceil(len(c) / 2.8)))) for c in out]

text = " ".join(
    " ".join("a%d" % i for i in range(20)) + "." for _ in range(5)
)
plan = chunks(text)
print(plan, sum(s for _, s in plan))

What this means for cost

Avatar videos are priced per second at a quality tier (standard, plus or max), so planned seconds are what you pay for. I am not listing a price here: check the cost on your own account for the exact figure for your script. What the plan shows is that the clip count, and the render time, change with sentence length even when the word count does not.

Steps

To get more words into each clip:

  • Count the words in each sentence.
  • Pair short sentences so that each pair adds up to 33 words or fewer.
  • Avoid a sentence of 34 words or more, which splits.
  • Re-run the plan and keep the total under 60 seconds.

What Sume does not do

Sume does not rewrite your script to fit the packing. The Avatar flow can plan a scene prompt for the first frame, but the words you submit are the words that are spoken. Editing for rhythm remains your job.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume