Split a 3-minute script into 60-second avatar jobs

Sume avatar talking videos accept scripts of an estimated 4-60 seconds. Split a 3-minute script into jobs of that size, then join the audio with timeline audio.

4 min readSume
All posts

A single Sume avatar talking-video job accepts a script only when Sume estimates the video at 4-60 seconds inclusive, so a 3-minute script has to become at least three jobs. Split at sentence boundaries, submit each part with its own Idempotency-Key, and join the results afterwards.

What the docs set as the limit

The avatar-video page states that scripts and multi-scene plans are accepted when Sume estimates the target duration at 4-60 seconds, and says to shorten longer scripts or split them into multiple jobs. The estimate is Sume's, not yours, so leave a margin when you split rather than cutting at a number you computed yourself.

The same window applies to video_inputs, and inline captions are rejected above 60 estimated seconds.

Avatar video request facts from the Sume docs (read 2026-10-03)
FieldValue
RoutePOST /v1/avatar-1.0/talking-video
Script or plan lengthEstimated 4-60 seconds, inclusive
AvatarTop-level avatar_handle, or per-scene character fields in video_inputs
Script inputExactly one of script or video_inputs
qualitystandard, plus (default) or max
resolution720p

Split on sentence boundaries

Cutting mid-sentence gives you two clips whose delivery does not match. The sketch below packs whole sentences into parts under a word budget you choose. The budget is your own conservative guess, not a Sume number: if a part comes back rejected for length, lower it and resubmit that part only.

import re

def split_script(text, max_words=110):
    sentences = re.split(r"(?<=[.!?])\s+", text.strip())
    parts, current = [], []
    for sentence in sentences:
        words = len(sentence.split())
        if current and sum(len(s.split()) for s in current) + words > max_words:
            parts.append(" ".join(current))
            current = []
        current.append(sentence)
    if current:
        parts.append(" ".join(current))
    return parts

script = "Welcome to the course. In this lesson we cover three steps. " * 20
for i, part in enumerate(split_script(script)):
    print(i, len(part.split()), "words")

Submit each part, then join the audio

Submit one job per part with the same avatar_handle and quality, and an Idempotency-Key such as lesson-7-part-1, so a retry returns the original job instead of billing twice. Do not resubmit a part because your client timed out: poll status_url instead.

Each finished avatar video is a separate MP4. If you want one continuous audio file, detach the audio from each clip and concatenate the parts with timeline audio. Its concat operation takes 1-20 ordered parts, all already on media.sume.com, and returns one audio_url plus the offsets of each segment. The result has no gaps and no re-synthesis.

For one continuous picture as well, place the clips in Timeline 1.0 and use the joined file as the audio track. Keep the same avatar and aspect ratio across parts so the cuts do not jump.

  • Same avatar_handle, quality and aspect_ratio on every part.
  • One Idempotency-Key per part, stable across retries.
  • Poll status_url; never resubmit a paid job for the same part.
  • Concat takes at most 20 parts per call.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume