Tavus Starter: 3 concurrent streams vs Sume bulk concurrency 1 to 16

Tavus Starter caps live conversations at 3 concurrent streams. Sume bulk runs queue up to 100 renders and keep 1 to 16 in flight. Different jobs.

5 min readSume
All posts

Tavus's pricing page lists up to 3 concurrent streams on Starter and up to 10 on Growth; a stream is a live video conversation, so the cap is how many people can be talking to a replica at once. Sume's bulk runs measure something else: you queue 1 to 100 Format runs and set concurrency from 1 to 16, which is how many renders stay in flight. A live stream occupies a seat while a person talks; a bulk window occupies a worker while a video renders, then frees it.

Tavus source: pricing, read 2026-10-04. Sume sources: Bulk runs and Runs and results.

Two meanings of concurrency

The word hides a real difference. For a live conversation, concurrency is demand you cannot defer: a third caller joins while two are talking, and the fourth is turned away or waits. For a render queue, demand can wait, and the only cost of a small window is time.

Concurrency compared (read 2026-10-04)
QuestionTavusSume bulk run
What is countedConcurrent live streamsChild Format runs in flight
Starter or entry limit3 streamsconcurrency 1 to 16, your choice
Larger tierGrowth: 10 streamsSame range; workspace generation concurrency still applies
What happens at the limitNot described on the page readRemaining items wait in the queue
Maximum batchNot applicable100 items per queue

Wave arithmetic for a batch

If each render takes about N minutes and your window is C, a queue of M items finishes in roughly M divided by C waves of N minutes. A 100-item queue at concurrency 10 is 10 waves. Sume's Format docs say long-form host video typically finishes in 15 to 30 minutes, so plan hours, not minutes, for a full batch of that kind.

  • Raise concurrency only if your workspace allows it; children still share workspace limits.
  • Each item can register its own communication.webhook_url.
  • The queue has no webhook and no cancel endpoint; cancel a child run instead.

When live is the requirement

If a viewer must talk back to the face, a rendered clip does not substitute. Tavus's product is built for that and Sume's avatar docs describe rendered, script-driven clips only. If the content is the same for every viewer, a rendered clip serves any number of viewers and has no stream cap at all.

Bottom line

Do not compare 3 streams with 16 workers as if they were the same number. Decide first whether you need conversation or a deliverable, then size the matching limit with your real traffic.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume