Tavus max_call_duration is plan-capped; Sume's cap is 60 s a job
Tavus ends a call at max_call_duration, capped by your plan; an unjoined call times out after 300 s. Sume's avatar video takes 4-60 s per job. Table inside.
On Tavus, a conversation ends when max_call_duration is reached, and the value you can set is capped by your subscription plan. A call nobody joins ends after participant_absent_timeout, which defaults to 300 seconds. On Sume the equivalent ceiling is not a call length but a render length: an avatar video script or multi-scene plan must estimate to 4-60 seconds per job.
The two limits answer different questions. Tavus bounds how long a live session can run. Sume bounds how long one finished clip can be.
The three Tavus timers
Tavus documents three parameters. max_call_duration sets the maximum call length in seconds and ends the conversation when reached, regardless of activity. If you ask for more than your plan allows, the value is silently capped to the plan maximum. participant_left_timeout (default 0 seconds) ends the call after the last participant leaves, and participant_absent_timeout (default 300 seconds) ends it if no one joins (Tavus docs: Call Duration and Timeout). Tavus points to the pricing page for the plan-specific maximum, so check that row before you promise a customer a 30 minute call.
| Question | Tavus conversation | Sume avatar video |
|---|---|---|
| What bounds the length | max_call_duration, capped by plan | Script or plan must estimate to 4-60 s |
| If you exceed it | Silently capped to the plan limit | Rejected; shorten or split into several jobs |
| Nobody shows up | Ends after participant_absent_timeout (300 s default) | Not applicable; a job renders without an audience |
| Unit of work | A live session | A job: queued, processing, completed |
Planning a longer piece on Sume
Sume's avatar docs say scripts and multi-scene plans are accepted when the estimated duration is 4-60 seconds inclusive, and longer scripts should be shortened or split into multiple jobs (Generate avatar video). A three minute explainer is therefore at least three jobs. Ordered pieces can then be assembled into one MP4 with Timeline 1.0, which takes one audio spine plus ordered video slots.
Split at paragraph boundaries and give each piece its own opening beat. Every piece is a separate render, so keep the same avatar handle and scene settings across them and check the joins on a finished file.
- Inline captions are rejected above 60 seconds of estimated duration, so caption each piece or caption the assembled file with the standalone captions model.
- Each piece is an independent paid job. A failed piece does not fail the others, and you can re-render only that piece.
Which limit to design around
If the experience is a conversation, design around the Tavus plan cap and the 300 second no-show timer, and tell users when time is nearly up. If the experience is a message, design around 60 seconds and treat length as a script-editing problem.
Questions to ask before you build
Before committing to either, write down the longest single piece of content a user will receive or take part in. For a live call, ask Tavus for the exact max_call_duration ceiling on your plan and test what the user sees at the cutoff. For a rendered piece, check that the longest script estimates to 60 seconds or less.
Also decide what happens at the limit. A call that ends should say so and offer a way to continue. A script that is too long should be split at a natural point, not truncated, so every piece stands alone.
Sources
Related posts
More in Comparisons
- Tavus max_participants counts the face: seats vs a shareable clip
In a Tavus conversation the AI face takes a seat: max_participants of 2 means one human plus one face. A Sume clip has no seats.
- Tavus screen share: the agent sees your screen; Sume takes images
Tavus screen share needs raven-1 perception, a live video room and a user who starts sharing. A Sume avatar clip takes a product image or scene photo instead.
- Tavus Sparrow-2 turn-taking vs writing pauses in a Sume clip
Tavus lets you tune turn_taking_patience and pal_interruptibility on a live agent. A Sume avatar clip has no turns: you author pauses as silence scenes.
- Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.
Written by Sume