Pipecat TavusVideoService live avatar, or a recorded Sume clip?
Pipecat's Tavus service speaks your agent's TTS live over WebRTC. A Sume avatar job is a 4 to 60 second recorded MP4. How to pick, by what the viewer does.
What does Pipecat's Tavus service do?
Pipecat's docs describe TavusVideoService as generating Tavus AI avatar video that speaks your Pipecat agent's TTS output in real time. It needs an api_key, a replica_id and an aiohttp.ClientSession, and it uses a dual-room setup with DailyTransport.
That is a live conversation component: the viewer talks, your agent answers, and the face moves as the audio streams. Sume does not ship a live full-duplex video session. Sume Avatar 1.0 renders a finished file.
How do the two compare?
| Pipecat TavusVideoService | Sume avatar job | |
|---|---|---|
| Output | Real-time video from your agent's TTS | MP4 file after a job completes |
| Interaction | Viewer and agent talk back and forth | None; script in, clip out |
| Length | A session | 4 to 60 seconds per job |
| Cost shape | Vendor session pricing | Per second: $0.184 to $0.58 by tier |
| Reuse | Each session is live | One clip plays for any number of viewers |
When is the recorded clip the better answer?
When every viewer should see the same thing: a product explainer, a holiday offer, an onboarding step. You render once, review the result and embed it. The cost does not grow with views, and you can approve the exact words before anyone sees them.
When the viewer has to ask something unscripted, a live service is the right tool. A common split is a recorded clip as the greeting and a live agent behind a button.
What does the recorded side cost?
A 30 second plus-tier clip with no product image is 30 x $0.245 = $7.35. At standard it is $5.52, and at max $16.50. Read the cost by tier for the product-image rates.
Jobs are async by default. For where live and rendered split, the earlier comparison goes through the Tavus conversation API.
Sources
Related posts
More in Sume Avatar 1.0
- Reuse one AI presenter across lip-sync clips with an avatar handle
Create the presenter once, then pass avatar_handle to H3 Max lip sync for every clip. Same face each time, no re-upload, one Idempotency-Key per line.
- Tavus Griffin-Lite is invite-only: what can you build today?
Griffin-Lite is a closed Tavus preview. Until you get in, Sume ships rendered avatar clips, lip sync and motion control for talking-head video.
- YouTube Shorts series: a weekly AI presenter episode pipeline
YouTube began rolling out Shorts series on 2026-09-23. A repeatable weekly pipeline for presenter episodes: one avatar, draft at standard, final at max.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume