Avatar script: 33 words = one 12 s clip, 34 words = two 7 s clips

A Sume Avatar 1.0 sentence of 33 words fits one 12-second clip. Add a 34th word and it splits into two 7-second clips. The numbers, with code.

4 min readSume
All posts

One more word can change how many clips your avatar video has. When a script is run through the Avatar 1.0 chunker, a single sentence of 33 words stays whole and is planned as one 12-second clip, while a sentence of 34 words is split into two balanced parts of 17 words, each planned as a 7-second clip. The reason is a ceiling of 33 words per clip, which is the 12-second maximum times 2.8 words per second, rounded down.

What the planner does

The numbers on this page come from running the avatar-workflows package in the Sume repo on constructed scripts, and from the Avatar videos docs for the 4 to 60 second window. The chunker is an internal planner: it decides how many clips your script becomes, and the public docs describe the outcome rather than the algorithm, so treat the exact constants as current behavior that can change.

The rules that matter here are small. Speech is estimated at 2.8 words per second. A clip is 4 to 12 seconds, so the longest clip holds floor(12 x 2.8) = 33 words. A sentence longer than that is cut into the fewest equal-sized parts that each fit. Duration is the word count divided by 2.8, rounded up, then clamped to 4 to 12.

One sentence, different lengths (read 2026-10-08)
Words in the sentenceClipsWords per clipPlanned seconds
101104
3313312
34217 and 177 and 7 (14 total)
50225 and 259 and 9 (18 total)

Why this matters for the window

Total length is the sum of clip durations, and the talking video must come in between 4 and 60 seconds. Because the split rounds up for each part, a long run-on sentence costs more planned time than the same words as short sentences. Two clips of 7 seconds are 14 seconds, while the words alone, 34 divided by 2.8, are about 12.1. The extra comes from the second ceiling. If you are close to the 60 second cap, end sentences earlier.

Mirror it before you submit

This small script reproduces the clip plan for text that uses full stops. It is a simplified mirror, not the production code, and it skips the rare tail merge for a last chunk of under 7 words.

import math, re
def chunks(script):
    maxw, out, cur = 33, [], []
    for s in re.findall(r"[^.!?]+[.!?]?", script):
        w = s.split()
        if len(w) > maxw:
            if cur:
                out.append(cur)
                cur = []
            n = math.ceil(len(w) / maxw)
            k, r = divmod(len(w), n)
            i = 0
            for j in range(n):
                size = k + (j < r)
                out.append(w[i:i + size])
                i += size
        elif cur and len(cur) + len(w) > maxw and len(cur) >= 7:
            out.append(cur)
            cur = w
        else:
            cur = cur + w
    if cur:
        out.append(cur)
    return [(len(c), min(12, max(4, math.ceil(len(c) / 2.8)))) for c in out]

def sentence(n):
    return " ".join("w%d" % i for i in range(n)) + "."

for n in (10, 33, 34, 50):
    print(n, chunks(sentence(n)))

Steps

Use the check as a gate in your pipeline.

  • Write the script in sentences that end with a full stop, question mark or exclamation mark.
  • Run the mirror and look for any sentence over 33 words; rewrite it as two sentences.
  • Add up the planned seconds and keep the total under 60.
  • Submit the script to POST /v1/avatar-1.0/talking-video with an Idempotency-Key.
  • Compare the chunk count the job reports with your plan.

What Sume does not do

Sume does not let you set where a clip starts or ends from a plain script. If you need exact cuts, use scenes (video_inputs) instead of a script; you cannot send both in one request. The word rate is also an estimate, so the real speech can run a little shorter or longer than the plan.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume