AI avatar video length limit: 4 to 60 seconds per job
Sume's avatar video accepts an estimated 4 to 60 seconds per job, for scripts, scenes, previews and inline captions. Where it applies and what to do outside it.
An avatar video job on Sume has to fit an estimated 4 to 60 seconds, inclusive. The rule is about Sume's estimate of the target duration, not a word or character count, and it applies to a plain script, to ordered video_inputs, to previews and to inline captions. Outside the window, shorten the script or split it into multiple jobs.
Where does the 4 to 60 second window apply?
| Surface | Rule |
|---|---|
script on talking-video | Estimated 4-60 seconds inclusive |
Multi-scene video_inputs | Total planned duration must land in 4-60 seconds |
| Avatar video preview | Same window: estimated 4-60 seconds inclusive |
Inline captions | Estimated duration above 60 seconds is rejected |
Is the limit a word count?
No. The docs say Sume estimates the target video duration, and they publish no formula, so do not treat any words-per-minute figure as a guarantee. For a rule of thumb and a worked example, see script length for an AI avatar video; this post only covers where the window applies.
In a multi-scene plan you set each scene's voice.duration yourself. The docs example uses scenes of 3, 4 and 5 seconds, which adds up to 12 seconds and sits inside the window.
Is there a minimum too?
Yes. The window is 4 to 60 seconds inclusive, so a one-word script whose estimate falls under 4 seconds is outside it as well. Silence beats count for this, because a scene with voice.type of silence carries its own duration.
The same window is what the preview route uses, so you can test a borderline script there before spending on a render.
What happens when my script is too long?
The docs tell you to shorten longer scripts or split them into multiple jobs. Each job is its own request with its own Idempotency-Key, and each returns a clip you join afterwards; the joining steps are in AI avatar video longer than 60 seconds.
Captions are the second place a long video bites. Inline captions reject an estimated duration above 60 seconds, but the standalone Video captions route takes an existing public video URL, so it is the route for a clip you have already joined. See captions on a long video.
How do I stay inside the window?
- Write the script first and cut it until it reads comfortably in under a minute.
- Prefer several short clips over one clip that sits near 60 seconds; a single estimate near the edge is the easiest to overshoot.
- Use a preview first: it applies the same window, so a script that is too long fails before you pay for the full render.
- For a scene that should stay quiet, use a silence beat with its own
durationrather than padding the script.
Sources
Related posts
More in Developers
- Talking photo AI without watermark: what Sume documents
Sume's docs describe no watermark option for talking photo clips. They do fix the inputs: a still, Sume-hosted audio under 10 MB, 1-300 s or 5-14.8 s by route.
- Team API keys: which wallet does a Sume run charge?
A Sume run charges the workspace its API key belongs to. A team key spends the team wallet; a personal key on a team Format is refused with 403.
- C# text to speech: call a TTS API with HttpClient
Text to speech in C#: POST the text and a voice with HttpClient, poll the job until it finishes, then stream the MP3 from its audio_url to a file.
- Text to speech streaming API: what Sume returns instead
Sume's text to speech API doesn't stream audio chunks. It returns a finished file per job; split long scripts into sentence jobs to start playback sooner.
Written by Sume