Tavus knowledge base limits vs putting the facts in a Sume script
Tavus knowledge bases are English-only, take 5-10 minutes to process and cap crawls at 10,000 documents. A Sume avatar script carries its facts in the text.

A Tavus knowledge base is for a live agent that must answer questions from your documents. It supports English documents only, takes 5-10 minutes to process, and caps website crawling at 10,000 crawl documents per account. A Sume avatar video has no retrieval step: the facts the avatar says are the words in the script you submit.
If your content is a fixed pitch, you do not need a knowledge base. If viewers ask unplanned questions, you do.
Tavus knowledge base limits
Tavus lists supported file types as .pdf, .txt, .docx, .doc, .png, .jpg, .pptx, .csv and .xlsx. Processing can take 5-10 minutes depending on document size. For crawls it lists a maximum of 10,000 crawl documents per account, 5 concurrent crawls, a 1-hour cooldown between recrawls of the same document, and max_pages of 1-100 per crawl. The page states no size limit for documents (Tavus docs: Knowledge Base).
| Item | Limit |
|---|---|
| Language | English documents only, works best for English conversations |
| Processing time | 5-10 minutes depending on size |
| Crawl documents per account | 10,000 |
| Concurrent crawls | 5 |
| Recrawl cooldown | 1 hour per document |
| Pages per crawl | max_pages 1-100 |
The scripted alternative
On Sume you submit the text. Nothing is retrieved or improvised, so a price, a date or a policy appears in the clip exactly as you wrote it. The cost is that the clip is only as current as the script: if a price changes, you render again (Generate avatar video).
For a personalised batch, pull the facts from your own database at the moment you build each script, then submit one job per recipient. Hold the source record next to the job id so you can prove what the clip said.
- Keep numbers that change out of long-lived clips, or put them in the page around the video.
- There is no retrieval layer, so the English-only knowledge base limit does not apply to a script; check the avatar docs for the languages your voice supports.
- A script over 60 seconds of estimated duration must be split into several jobs.
Decision rule
Ask whether you can list every question the viewer will have. If yes, script them and render. If no, a live agent with a knowledge base is the right tool, and the Tavus limits above are the ones to plan for.
Planning around the 5-10 minutes
A knowledge base is not instant, so a flow that creates one at signup needs a waiting state or a fallback. Process documents ahead of time and attach them when the conversation is created.
For crawls, remember the cooldown of one hour before the same document can be recrawled, so a page that changes hourly will lag. In a scripted flow there is no such lag: the script is built when you submit.
Keeping scripted facts correct
Scripts go stale in a way retrieval does not. Give each clip an owner and an expiry date in your own records, and re-render when a fact changes. Store the script text with the job id so you can find every clip that contains an old number.
For facts that change often, keep them out of the video and put them in text beside it, where an update is a page edit rather than a new render.
Sources
Related posts
More in Comparisons
- Tavus Magic Canvas cards vs what a rendered Sume avatar clip shows
Tavus Magic Canvas shows 8 kinds of interactive cards in live video calls only. A Sume avatar clip is a fixed MP4: CTA goes in a closing scene.
- Tavus max_call_duration is plan-capped; Sume's cap is 60 s a job
Tavus ends a call at max_call_duration, capped by your plan; an unjoined call times out after 300 s. Sume's avatar video takes 4-60 s per job. Table inside.
- Tavus max_participants counts the face: seats vs a shareable clip
In a Tavus conversation the AI face takes a seat: max_participants of 2 means one human plus one face. A Sume clip has no seats.
- Tavus screen share: the agent sees your screen; Sume takes images
Tavus screen share needs raven-1 perception, a live video room and a user who starts sharing. A Sume avatar clip takes a product image or scene photo instead.
Written by Sume