What a scripted AI avatar clip cannot do: seven limits and fixes

A Sume avatar clip cannot take questions, run past 60 seconds or switch presenter mid-video. Seven documented limits, each with a workaround.

5 min readSume
All posts

A Sume avatar video is a rendered file of a presenter reading your script. It cannot hold a conversation, react to the viewer, run longer than a minute in one job, or use two presenters in one final video. Knowing these limits before you design a campaign saves a rewrite later. Each is stated in the docs or the product, and each has a workaround.

The seven limits

Limits of Sume avatar clips and workarounds (Sume docs, read 2026-10-07)
LimitDetailWorkaround
No live conversationThe route renders a job; there is no real-time sessionPre-render answers; use a live agent product for open questions
4 to 60 secondsEstimated duration outside the window is rejectedSplit the script into several jobs
One avatar per final videoExecution supports one resolved avatar and a shared sceneRender one clip per speaker and cut them together elsewhere
720p onlyResolution is 720p at this timePlan for phones and small screens; do not promise 1080p
English onlyAvatar 1.0 speaks EnglishAdd subtitle cues in other languages; record a human voice if the message is critical
Face swap is shortBeta source video is about 4 to 15 seconds with usable audioCut the source clip first
Captions come from the spoken textInline captions never appear on preview stillsAdd authored cues with the standalone captions job

Why some limits are worth keeping in mind

The one-avatar rule matters for dialogue. A two-person conversation is two clips, one per speaker, cut together by you. The same rule explains why multi-scene video_inputs share one avatar and one background: scenes are beats in the same setting, not different locations.

The 60-second window also applies to inline captions, since Sume rejects an estimated duration above 60 seconds for them.

The big one: no conversation

The clip says what you scripted, once. It cannot interrupt, wait for a reply or notice a viewer's face. If your design needs any of that, the product category you want is a live conversational agent, which Sume does not provide. A scripted avatar can still carry the high-volume part of the job: the ten answers customers ask most, the welcome video, the explainer, the recap.

  • Use a clip when the message is fixed and you want to review it first.
  • Use a live agent when the viewer's questions are unpredictable.
  • Use both when you can: a clip first, with a link to a human or a live agent after it.

What to put in the page around the clip

State that the video is AI-generated. Offer a text version. Give a human contact route. These three lines cover the cases where the clip's limits matter most to the viewer.

How to test a limit before you commit

Create a preview with your real script, check that the estimate passes the 4 to 60 second check, and look at the stills for the scene you need. Previews are cheap compared with a full render, and they surface length and composition problems early. For language, listen to a short test clip with your actual wording, including names and acronyms.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume