What Sume Avatar 1.0 does not do: eight limits to check first

No streaming, no interruption, English-only speech in code, 720p, 4 to 60 seconds, one avatar per video. The limits of Sume Avatar 1.0 in one table.

5 min readSume
All posts

Sume Avatar 1.0 makes finished, script-driven talking videos. It does not stream, listen or interrupt, its speech prompt is English only in code, and every video is 720p, 4 to 60 seconds, with one avatar. Check these eight limits against your use case before you design around it; none of them is a defect, they are what the product is.

Avatar 1.0 limits, from the Sume docs and code on main as of 2026-10-08
LimitWhat is trueWhere it comes from
Not real timeJobs are queued, processing, then completedJobs and results docs
No listeningInput is a script or scene plan, not live audioGenerate avatar video docs
English speechClip prompt: English only, no non-English speechWorkflow code on main
LengthEstimated 4 to 60 secondsDocs
Resolution720p at this timeDocs
One avatar per videoOne resolved avatar and one shared sceneDocs
Face swap is BetaSource about 4 to 15 s, no promptsFace swap docs
WebhooksTerminal events only, no progressWebhooks docs

The real-time gap

Streaming models exist. Tavus describes Griffin-Lite as a full-duplex, real-time model but says it is not available to customers and is open to select trusted testers. Whether you can use it is the vendor's decision. Sume does not list a model with that behavior, and Sume's documents do not describe a live session endpoint.

What you can do now is cover the questions you can predict with rendered clips, covered in the Sume Avatar 1.0 docs.

The workarounds that are real

Some limits have a documented way around them:

  • Over 60 seconds: split the script into separate jobs, or use scenes if the total stays inside the window.
  • Different language: use TTS with a language field and a lip-sync route, and test the result.
  • Progress bars: poll the job events endpoint, since webhooks carry terminal events only.
  • Review before spend: first-frame previews, then generate-video.

The workarounds that are not

Do not chain clips to fake a conversation and call it live. The viewer waits for each render, and the cost is per clip. Do not describe a rendered avatar as live or as listening in your own copy. If the use case needs a real-time answer, the honest options today are a person, text chat, or a vendor with a live product that you have access to.

Ask the vendor the same questions

Use the table as a checklist for any avatar product, not only Sume. Ask about latency, language, length, resolution, number of people on screen, review before spend, and what happens when a job fails. A vendor that answers each in writing, with a date, is easier to build on than one that answers in a launch post.

For Tavus Griffin-Lite, the page we read on 2026-10-08 gives availability, latency and study figures, and says nothing about language, so that question is still open.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume