Tavus CVI has six layers; a Sume avatar video is one job
Tavus CVI chains perception, turn-taking, speech, an LLM, TTS and rendering. A Sume talking-video is one request with a script. Which parts you own in each.
What Tavus says CVI is
Tavus describes CVI as a PAL, a Face (Phoenix) and a Conversation over WebRTC. Under that sit six layers: Raven for visual perception, Sparrow for conversational flow, speech recognition, an LLM, text to speech and Phoenix avatar rendering (Tavus CVI docs, read 2026-10-05).
That is a real-time stack. A viewer speaks, the system listens, decides when to answer and renders a face on a live stream.
What a Sume avatar video asks of you
You send a script (or a video_inputs plan), an avatar_handle, a quality tier and an aspect ratio. Sume returns a job. The job moves queued, processing, then completed, failed or canceled. There is no listening layer, no live turn-taking and no session to hold open.
| Layer | Tavus CVI | Sume talking-video |
|---|---|---|
| Listening and perception | Raven, speech recognition | Not part of the request |
| Deciding what to say | LLM and Sparrow flow | You write the script |
| Voice | TTS layer | Avatar voice, must be ready |
| Face rendering | Phoenix | One rendered MP4 per job |
| Transport | WebRTC conversation | Poll, or a webhook on job.completed |
A minimal request
One call replaces the stack when the words are known in advance.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: cvi-compare-001" \
-d '{"avatar_handle":"product_host","quality":"standard","script":"Three things changed in this release, and the first saves you an hour a week."}'Choosing
If the viewer must interrupt and be answered in real time, you need a conversation product. If you know the words, a clip job is simpler to test, cache and re-run. A 10 s standard clip is $1.84.
Sources
Related posts
More in Comparisons
- Tavus Starter $59 for 100 minutes vs 100 minutes of Sume avatar video
Tavus Starter lists 100 minutes at $59 a month. Here is what 100 minutes of finished Sume avatar video costs at standard, plus and max, and why it is not equal.
- 10-minute narrated lesson from slides: Synthesia Basic vs Sume
Synthesia's free Basic plan lists 500 credits, about 10 minutes. A narrated slide lesson on Sume is about $1.43 for 10 minutes, with no presenter.
- Test a 'noisy audio' transcription claim: 5 clips for 5 cents on Sume
Microsoft says MAI-Transcribe-2 handles noisy audio. Run five 60-second noisy clips through Sume STT for $0.05 and judge the text yourself.
- Text-to-speech price per million characters: 11 rates vs Sume
Eleven TTS rates converted to dollars per 1M characters, from $11 to $150. Sume TTS 1.0 is $47.50. Each is read from the vendor page, and caveats are listed.
Written by Sume