What Sume Avatar 1.0 does not do: eight limits to check first
No streaming, no interruption, English-only speech in code, 720p, 4 to 60 seconds, one avatar per video. The limits of Sume Avatar 1.0 in one table.
Sume Avatar 1.0 makes finished, script-driven talking videos. It does not stream, listen or interrupt, its speech prompt is English only in code, and every video is 720p, 4 to 60 seconds, with one avatar. Check these eight limits against your use case before you design around it; none of them is a defect, they are what the product is.
| Limit | What is true | Where it comes from |
|---|---|---|
| Not real time | Jobs are queued, processing, then completed | Jobs and results docs |
| No listening | Input is a script or scene plan, not live audio | Generate avatar video docs |
| English speech | Clip prompt: English only, no non-English speech | Workflow code on main |
| Length | Estimated 4 to 60 seconds | Docs |
| Resolution | 720p at this time | Docs |
| One avatar per video | One resolved avatar and one shared scene | Docs |
| Face swap is Beta | Source about 4 to 15 s, no prompts | Face swap docs |
| Webhooks | Terminal events only, no progress | Webhooks docs |
The real-time gap
Streaming models exist. Tavus describes Griffin-Lite as a full-duplex, real-time model but says it is not available to customers and is open to select trusted testers. Whether you can use it is the vendor's decision. Sume does not list a model with that behavior, and Sume's documents do not describe a live session endpoint.
What you can do now is cover the questions you can predict with rendered clips, covered in the Sume Avatar 1.0 docs.
The workarounds that are real
Some limits have a documented way around them:
- Over 60 seconds: split the script into separate jobs, or use scenes if the total stays inside the window.
- Different language: use TTS with a language field and a lip-sync route, and test the result.
- Progress bars: poll the job events endpoint, since webhooks carry terminal events only.
- Review before spend: first-frame previews, then generate-video.
The workarounds that are not
Do not chain clips to fake a conversation and call it live. The viewer waits for each render, and the cost is per clip. Do not describe a rendered avatar as live or as listening in your own copy. If the use case needs a real-time answer, the honest options today are a person, text chat, or a vendor with a live product that you have access to.
Ask the vendor the same questions
Use the table as a checklist for any avatar product, not only Sume. Ask about latency, language, length, resolution, number of people on screen, review before spend, and what happens when a job fails. A vendor that answers each in writing, with a date, is easier to build on than one that answers in a launch post.
For Tavus Griffin-Lite, the page we read on 2026-10-08 gives availability, latency and study figures, and says nothing about language, so that question is still open.
Sources
Related posts
More in Sume Avatar 1.0
- YouTube AI disclosure: which avatar pipeline steps are exempt?
YouTube exempts scripts, captions and upscaling from its AI label but not realistic synthetic people. Map each step of a Sume avatar pipeline to the rule.
- Does a scripted AI presenter video need YouTube's synthetic label?
YouTube asks for a label when content makes a real person appear to say something they did not. Here is how to read that for a made-up Sume Avatar presenter.
- Will a YouTube AI label hurt reach? What the page says
YouTube's help page says disclosing AI content does not limit reach or monetization eligibility; penalties target non-disclosure. Plan an avatar series on that.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume