Tavus Sparrow-2 turn-taking vs writing pauses in a Sume clip
Tavus lets you tune turn_taking_patience and pal_interruptibility on a live agent. A Sume avatar clip has no turns: you author pauses as silence scenes.
Tavus exposes turn-taking as settings on a live agent: turn_detection_model (sparrow-1 or sparrow-2), turn_taking_patience and pal_interruptibility, each low, medium or high. A Sume avatar clip has no turns to manage, so pacing is authored: a pause is a video_inputs scene with voice.type: "silence" and a duration you choose.
If you are comparing the two for a sales or support flow, the question is whether the viewer will talk back. If they will, you need turn-taking. If they will not, you need good pacing.
The Tavus settings
Tavus lists sparrow-1 and sparrow-2 as the turn detection models. The page gives sparrow-1 as the default and says the default changes to sparrow-2 on September 8, 2026; it calls Sparrow-2 the frontier model and lets you opt in by setting turn_detection_model. turn_taking_patience runs from low (quick replies that may cut in on natural pauses) to high (waits for a clear end of turn), with medium as the default. pal_interruptibility sets how readily the agent stops when the user speaks (Tavus docs: Conversational Flow). Because the page ties the default to a date, check it for the model your account gets today.
| Setting | Values | Default |
|---|---|---|
turn_detection_model | sparrow-1, sparrow-2 | sparrow-1; the page says it changes to sparrow-2 on Sept 8, 2026 |
turn_taking_patience | low, medium, high | medium |
pal_interruptibility | low, medium, high | medium |
idle_engagement | off, patient, eager | off |
Pacing in a rendered clip
In a Sume multi-scene avatar video each entry in video_inputs has a voice. A spoken scene uses type: "text" with a script; a non-speaking beat uses type: "silence" with a required duration and no script (Generate avatar video). The total planned duration must still land inside the 4-60 second window.
Write the pause where a live agent would wait for an answer: after a question to the viewer, before a price, after a surprising number. A two second silence scene reads as the agent giving the viewer time to think.
- Keep pauses short. In a 30 second clip, one or two beats of 1-3 seconds is usually enough.
- Silence scenes count toward the planned duration, so a script near the 60 second limit has less room for beats.
- Captions are built from the spoken text, so a silence scene adds a gap rather than words.
Which approach fits
Pick live turn-taking when users reply. Pick authored pacing when you control the whole message. Griffin, which Tavus describes as a full-duplex model that continuously decides when to respond, is offered only as a restricted Griffin-Lite research preview for select testers, and Tavus says it is not available to customers at this time, so today's live option is the Sparrow-based pipeline.
A worked pacing example
Take a 30 second clip that asks a question. Write a spoken scene of about 12 seconds that sets up the problem, a 2 second silence scene, a spoken scene of about 10 seconds with the answer, a 1 second silence, and a closing line of about 5 seconds. The total stays under the 60 second ceiling with room to spare.
Preview the result before a long run. Because beats are authored, you can change a duration and render again rather than retuning a model setting, and the change is exactly what you asked for.
Sources
Related posts
More in Comparisons
- Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.
- Together AI TTS runs $4 to $65 per million characters; Sume is $47.50
Together AI lists text-to-speech from $4 to $65 per million characters. Sume TTS is $0.0475 per 1,000, or $47.50 per million. Cost of 100 scripts on each.
- Transcription cost: Scribe v2 $0.22/hour vs Sume's $0.01 per minute
ElevenLabs Scribe v2 lists $0.22 per hour of audio; Sume's video inspect transcript is $0.01 per audio minute ($0.60 per hour). Cost of 10, 60 and 600 minutes.
- Two TTS vendors both claim number one: how to settle it yourself
ElevenLabs and Cartesia each cite a top Artificial Analysis spot in September 2026. How to read both claims and run a blind test on Sume with Sonic model ids.
Written by Sume