What a scripted AI avatar clip cannot do: seven limits and fixes
A Sume avatar clip cannot take questions, run past 60 seconds or switch presenter mid-video. Seven documented limits, each with a workaround.
A Sume avatar video is a rendered file of a presenter reading your script. It cannot hold a conversation, react to the viewer, run longer than a minute in one job, or use two presenters in one final video. Knowing these limits before you design a campaign saves a rewrite later. Each is stated in the docs or the product, and each has a workaround.
The seven limits
| Limit | Detail | Workaround |
|---|---|---|
| No live conversation | The route renders a job; there is no real-time session | Pre-render answers; use a live agent product for open questions |
| 4 to 60 seconds | Estimated duration outside the window is rejected | Split the script into several jobs |
| One avatar per final video | Execution supports one resolved avatar and a shared scene | Render one clip per speaker and cut them together elsewhere |
| 720p only | Resolution is 720p at this time | Plan for phones and small screens; do not promise 1080p |
| English only | Avatar 1.0 speaks English | Add subtitle cues in other languages; record a human voice if the message is critical |
| Face swap is short | Beta source video is about 4 to 15 seconds with usable audio | Cut the source clip first |
| Captions come from the spoken text | Inline captions never appear on preview stills | Add authored cues with the standalone captions job |
Why some limits are worth keeping in mind
The one-avatar rule matters for dialogue. A two-person conversation is two clips, one per speaker, cut together by you. The same rule explains why multi-scene video_inputs share one avatar and one background: scenes are beats in the same setting, not different locations.
The 60-second window also applies to inline captions, since Sume rejects an estimated duration above 60 seconds for them.
The big one: no conversation
The clip says what you scripted, once. It cannot interrupt, wait for a reply or notice a viewer's face. If your design needs any of that, the product category you want is a live conversational agent, which Sume does not provide. A scripted avatar can still carry the high-volume part of the job: the ten answers customers ask most, the welcome video, the explainer, the recap.
- Use a clip when the message is fixed and you want to review it first.
- Use a live agent when the viewer's questions are unpredictable.
- Use both when you can: a clip first, with a link to a human or a live agent after it.
What to put in the page around the clip
State that the video is AI-generated. Offer a text version. Give a human contact route. These three lines cover the cases where the clip's limits matter most to the viewer.
How to test a limit before you commit
Create a preview with your real script, check that the estimate passes the 4 to 60 second check, and look at the stills for the scene you need. Previews are cheap compared with a full render, and they surface length and composition problems early. For language, listen to a short test clip with your actual wording, including names and acronyms.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume