Avatar script over 60 seconds: split it into jobs and join the clips
Sume Avatar 1.0 accepts 4 to 60 seconds per job. For a longer explainer, split the script at sentence breaks, send one job per part and keep the order.
Sume rejects an avatar script whose estimated duration is outside 4 to 60 seconds, and the docs say what to do: shorten it or split it into multiple jobs. For a three-minute talking-head explainer that means about four to six jobs, one per chapter, each with its own Idempotency-Key, rendered with the same avatar_handle, quality and aspect_ratio so the parts match.
Sume gives you the parts; stitching them into one file is your editor's job. Sume's Avatar route does not concatenate jobs.
Plan the split yourself
Split at sentence ends, not mid-sentence, and keep each part well under the limit. The words-per-second figure below is a planning assumption of this example, not a Sume number; the real check is Sume's own duration estimate, so a part can still be rejected.
import re
WPS = 2.2 # planning assumption, not a Sume figure
MAX_SECONDS = 50 # leave headroom under the 60 s limit
def split_script(text):
parts, cur = [], []
for s in re.split(r"(?<=[.!?])\s+", text.strip()):
trial = " ".join(cur + [s])
if cur and len(trial.split()) / WPS > MAX_SECONDS:
parts.append(" ".join(cur))
cur = [s]
else:
cur.append(s)
if cur:
parts.append(" ".join(cur))
return parts
parts = split_script("One. " * 300)
print(len(parts), [len(p.split()) for p in parts])Keep the parts consistent
- Use one
avatar_handle, onequality, oneaspect_ratiofor every part. - Key each job by part:
explainer:v1:part-01,part-02, and so on. A retry of part 3 then cannot create a second part 3. - Submit all parts, then poll each
GET /v1/jobs/{id}/status. The results are separatemedia.sume.comvideo artifacts. - If a part is under 4 seconds it is rejected, so merge a one-line ending into the part before it.
A table of the limits
| Item | Limit |
|---|---|
| Duration, script or video_inputs | 4 to 60 seconds, inclusive |
| Resolution | 720p |
| Inline captions | Rejected when the estimated duration is above 60 seconds |
| Scripts per request | Exactly one of script or video_inputs |
What Sume does not do
There is no "long video" mode for Avatar 1.0 in the docs, and no join endpoint for the parts. Join the finished parts in your own editor, or with a Timeline render if your workflow already uses one.
Sources
Related posts
More in Sume Avatar 1.0
- Write an avatar script that sounds authentic in 4 to 60 seconds
HeyGen says 64.6% trust avatars that sound authentic. Draft a Sume avatar script in the 4-60 second window with scene beats and a silence gap.
- Avatar video aspect ratios: 9:16, 1:1, 4:3 or 16:9 for ads?
Sume avatar videos render in 1:1, 3:4, 9:16, 4:3 or 16:9 at 720p. Which ratio to pick for feed, story and listing placements, and what the default is.
- Make an avatar video look professional: scene photo of your office
53.9% of shelved videos were held back for not looking professional enough, per HeyGen. Set a Sume avatar scene from your own photo, prompt or per-scene image.
- Avatar preview: regenerate stills or create a new preview
Regenerate redraws the first-frame stills of one preview. A new script, avatar, scene or aspect ratio needs a new preview. Quality is set at generate-video.
Written by Sume