Tavus max_participants counts the face: seats vs a shareable clip
In a Tavus conversation the AI face takes a seat: max_participants of 2 means one human plus one face. A Sume clip has no seats.

In a Tavus conversation the face counts as a participant, so max_participants: 2 allows one human plus one face. A third viewer cannot join once the limit is reached. A Sume avatar video has no seats: it is a finished MP4, and the limit you plan around is how many render jobs your workspace can run, not how many people watch.
This matters when the plan is a demo for a group: a live session is one conversation per room, while a rendered clip is a file you can send to everyone.
What the Tavus parameter says
Tavus's participant-limits page defines max_participants as the control on how many participants are allowed in a conversation, notes that faces are counted, and says that when the limit is reached additional users cannot join. The page does not state a default or a numeric maximum (Tavus docs: Participant Limits). If you need an audience larger than one, ask Tavus what your plan allows before building the flow.
| Setting | Tavus | Sume avatar video |
|---|---|---|
| Who occupies a seat | Each human and each face | No seats; a clip is a file |
| Typical one-on-one value | max_participants: 2 | Not applicable |
| What limits throughput | Plan and participant settings | Workspace concurrency and queue capacity |
| Viewers of the output | Those in the room | Anyone you share the URL with |
What limits Sume instead
Sume limits paid generation by processing concurrency and queue capacity: Free runs 1 and queues 5, Pro runs 4 and queues 20, Startup 8 and 40, Scale 20 and 100. A full processing limit does not reject submits; valid jobs wait as queued until a slot opens, and only a full queue returns 429 queue_full (Generation admission).
So a campaign that personalises one clip for 200 recipients is a 200-job plan, paced by the queue, and the audience for each finished clip is whatever you decide.
- Live and rendered can be combined: a live session for the first call, a rendered recap clip for the people who were not on it.
- Treat the media URL as shareable: store the
media.sume.comURL and serve or embed it yourself.
Choosing by audience size
One viewer who needs to ask questions is a live conversation. Many viewers who need the same message is a rendered clip. In between, such as a small group meeting, the number of seats matters and you should confirm it with Tavus.
For rendered clips the cost is per job, not per viewer, so reuse is free. Generate once, host the file and share the URL. Only personalise per recipient when the personalisation changes the words, because each variant is a separate paid render.
Whichever you pick, test the experience for the last allowed viewer, not only for the first.
A planning example
Say a team wants 50 prospects to see a personalised greeting. On a live platform that is 50 conversations, each with a seat for the prospect and one for the face. On Sume it is 50 render jobs, queued as capacity allows, and 50 files that each prospect can open at any time afterwards.
The second approach moves the bottleneck from attendance to rendering, which you can schedule overnight. The first gives richer interaction but needs the prospect to be present.
Sources
Related posts
More in Comparisons
- Tavus screen share: the agent sees your screen; Sume takes images
Tavus screen share needs raven-1 perception, a live video room and a user who starts sharing. A Sume avatar clip takes a product image or scene photo instead.
- Tavus Sparrow-2 turn-taking vs writing pauses in a Sume clip
Tavus lets you tune turn_taking_patience and pal_interruptibility on a live agent. A Sume avatar clip has no turns: you author pauses as silence scenes.
- Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.
- Together AI TTS runs $4 to $65 per million characters; Sume is $47.50
Together AI lists text-to-speech from $4 to $65 per million characters. Sume TTS is $0.0475 per 1,000, or $47.50 per million. Cost of 100 scripts on each.
Written by Sume