Split a 3-minute script into 60-second avatar jobs
Sume avatar talking videos accept scripts of an estimated 4-60 seconds. Split a 3-minute script into jobs of that size, then join the audio with timeline audio.
A single Sume avatar talking-video job accepts a script only when Sume estimates the video at 4-60 seconds inclusive, so a 3-minute script has to become at least three jobs. Split at sentence boundaries, submit each part with its own Idempotency-Key, and join the results afterwards.
What the docs set as the limit
The avatar-video page states that scripts and multi-scene plans are accepted when Sume estimates the target duration at 4-60 seconds, and says to shorten longer scripts or split them into multiple jobs. The estimate is Sume's, not yours, so leave a margin when you split rather than cutting at a number you computed yourself.
The same window applies to video_inputs, and inline captions are rejected above 60 estimated seconds.
| Field | Value |
|---|---|
| Route | POST /v1/avatar-1.0/talking-video |
| Script or plan length | Estimated 4-60 seconds, inclusive |
| Avatar | Top-level avatar_handle, or per-scene character fields in video_inputs |
| Script input | Exactly one of script or video_inputs |
| quality | standard, plus (default) or max |
| resolution | 720p |
Split on sentence boundaries
Cutting mid-sentence gives you two clips whose delivery does not match. The sketch below packs whole sentences into parts under a word budget you choose. The budget is your own conservative guess, not a Sume number: if a part comes back rejected for length, lower it and resubmit that part only.
import re
def split_script(text, max_words=110):
sentences = re.split(r"(?<=[.!?])\s+", text.strip())
parts, current = [], []
for sentence in sentences:
words = len(sentence.split())
if current and sum(len(s.split()) for s in current) + words > max_words:
parts.append(" ".join(current))
current = []
current.append(sentence)
if current:
parts.append(" ".join(current))
return parts
script = "Welcome to the course. In this lesson we cover three steps. " * 20
for i, part in enumerate(split_script(script)):
print(i, len(part.split()), "words")Submit each part, then join the audio
Submit one job per part with the same avatar_handle and quality, and an Idempotency-Key such as lesson-7-part-1, so a retry returns the original job instead of billing twice. Do not resubmit a part because your client timed out: poll status_url instead.
Each finished avatar video is a separate MP4. If you want one continuous audio file, detach the audio from each clip and concatenate the parts with timeline audio. Its concat operation takes 1-20 ordered parts, all already on media.sume.com, and returns one audio_url plus the offsets of each segment. The result has no gaps and no re-synthesis.
For one continuous picture as well, place the clips in Timeline 1.0 and use the joined file as the audio track. Keep the same avatar and aspect ratio across parts so the cuts do not jump.
- Same avatar_handle, quality and aspect_ratio on every part.
- One Idempotency-Key per part, stable across retries.
- Poll status_url; never resubmit a paid job for the same part.
- Concat takes at most 20 parts per call.
Sources
Related posts
More in Sume Avatar 1.0
- Udemy promo video at 90 s: an avatar intro in two clips
Udemy's guidance puts an ideal promo video near 90 seconds. A Sume avatar clip tops out at 60 seconds, so build the intro as two clips and join them.
- Alt text for an avatar video poster: WCAG 1.1.1 in practice
A poster image from a Sume avatar preview still needs a text alternative under WCAG 1.1.1, and the video needs descriptive identification. What to write.
- Silent avatar video clip: what WCAG 1.2.1 asks for
A silent beat or silent clip in an avatar video is prerecorded video-only content. Here is what WCAG 1.2.1 wants and how to supply it from Sume.
- WCAG 1.2.3 for an AI avatar video: is the script enough?
A talking-head avatar video already has its words in a script. Here is when that text meets WCAG 1.2.3 and what to add if the picture carries more.
Written by Sume