Tavus Starter: 3 concurrent streams vs Sume bulk concurrency 1 to 16
Tavus Starter caps live conversations at 3 concurrent streams. Sume bulk runs queue up to 100 renders and keep 1 to 16 in flight. Different jobs.

Tavus's pricing page lists up to 3 concurrent streams on Starter and up to 10 on Growth; a stream is a live video conversation, so the cap is how many people can be talking to a replica at once. Sume's bulk runs measure something else: you queue 1 to 100 Format runs and set concurrency from 1 to 16, which is how many renders stay in flight. A live stream occupies a seat while a person talks; a bulk window occupies a worker while a video renders, then frees it.
Tavus source: pricing, read 2026-10-04. Sume sources: Bulk runs and Runs and results.
Two meanings of concurrency
The word hides a real difference. For a live conversation, concurrency is demand you cannot defer: a third caller joins while two are talking, and the fourth is turned away or waits. For a render queue, demand can wait, and the only cost of a small window is time.
| Question | Tavus | Sume bulk run |
|---|---|---|
| What is counted | Concurrent live streams | Child Format runs in flight |
| Starter or entry limit | 3 streams | concurrency 1 to 16, your choice |
| Larger tier | Growth: 10 streams | Same range; workspace generation concurrency still applies |
| What happens at the limit | Not described on the page read | Remaining items wait in the queue |
| Maximum batch | Not applicable | 100 items per queue |
Wave arithmetic for a batch
If each render takes about N minutes and your window is C, a queue of M items finishes in roughly M divided by C waves of N minutes. A 100-item queue at concurrency 10 is 10 waves. Sume's Format docs say long-form host video typically finishes in 15 to 30 minutes, so plan hours, not minutes, for a full batch of that kind.
- Raise
concurrencyonly if your workspace allows it; children still share workspace limits. - Each item can register its own
communication.webhook_url. - The queue has no webhook and no cancel endpoint; cancel a child run instead.
When live is the requirement
If a viewer must talk back to the face, a rendered clip does not substitute. Tavus's product is built for that and Sume's avatar docs describe rendered, script-driven clips only. If the content is the same for every viewer, a rendered clip serves any number of viewers and has no stream cap at all.
Bottom line
Do not compare 3 streams with 16 workers as if they were the same number. Decide first whether you need conversation or a deliverable, then size the matching limit with your real traffic.
Sources
Related posts
More in Comparisons
- TTS language counts in October 2026: Voxtral 9, MAI 23, ElevenLabs 90+
How many languages each TTS vendor states: Mistral Voxtral, Microsoft MAI-Voice-2.1, OpenAI and ElevenLabs, from vendor pages, plus Sume's language field.
- TTS latency numbers side by side: Eleven v4 Turbo, Voxtral, MAI Flash
Three vendors, three latency figures, three different things measured. A table of what each page says, and why a Sume TTS job is a different question.
- TTS leaderboard: 33 Elo points rank 5 to 12
On Versely's September 2026 voice leaderboard, ranks 5 to 12 span 33 Elo points. Here is what that gap means and how to test a voice with Sume's tts_create.
- Udio downloads are off: export-ready music for video work
Udio disabled downloads after its UMG settlement. Where to get an exportable AI music bed for video instead: Sume's Music Router returns a file URL.
Written by Sume