Talking head video vs conversational avatar: three things you get back
A talking-head video is a file; a conversational avatar is a session. Three differences, with Sume Avatar 1.0 and Tavus Griffin-Lite as examples.
A talking-head video is a finished file you can review and reuse, while a conversational avatar is a live session that reacts to a person. Searching for either term usually means one of three needs: a recording to publish, a face to talk to, or both. Sume Avatar 1.0 makes the first. The second is what Tavus describes for Griffin-Lite, which is limited to select trusted testers as of 2026-10-08.
| Question | Talking-head video (Sume Avatar 1.0) | Conversational avatar (Tavus Griffin-Lite, read 2026-10-08) |
|---|---|---|
| What you get back | A durable MP4 on a result URL | A real-time video stream |
| Input | Script or scene plan, avatar handle, quality | Reference image with streaming audio and controls |
| Who decides the words | You, before rendering | The conversation, as it happens |
| Reviewable before viewers see it | Yes: preview stills and the finished file | Not in advance |
| Access | Public API | Select trusted testers only |
1. A file versus a stream
A Sume job ends with a video_url on media.sume.com. You can open it, trim it, caption it and play it for a thousand viewers. A streaming model produces frames for one conversation, and there is no file unless you record the session yourself.
That changes storage and sharing. With a file you can send a link, and a retry is a new job with its own Idempotency-Key.
2. Review happens before or after
With a script you can read every word before spending money, and you can check first-frame stills with the preview endpoints. With a live model you review afterwards or you set rules in advance.
Tavus says Griffin-Lite can deceive a person into thinking it is not AI and says it is working on safe disclosure features. For a rendered clip, you control disclosure yourself: a caption cue, a line in the description, or a spoken sentence.
3. How you learn it finished
For a rendered clip the contract is a job. Sume sends signed job.completed, job.failed and job.canceled webhooks, terminal events only, and the docs say to keep polling in place as a fallback. There are no progress events in webhooks.
For a live session the question is not completion but continuity, which is a different engineering problem.
Picking one
If the viewer does not need to talk back, render a clip. If they do and you have access to a streaming model, use it for that part only and keep rendered clips for answers you already know.
A quick test for your own brief
Read your brief and ask whether the words can be written before the viewer arrives. If yes, you are describing a talking-head video, and Sume Avatar 1.0 can make it today at $0.184, $0.245 or $0.55 per second by tier. If the answer depends on what the viewer says, you are describing a conversational avatar, and the access question for any streaming vendor comes first.
Many briefs are a mix. Split them. The welcome, the product explanation and the sign-off are scripts. The question-and-answer part is a conversation. Treat them as two projects with two budgets.
Sources
Related posts
More in Comparisons
- Talking presenter: Fabric still plus voice, or video model plus voice?
Sume docs: video models do not lip-sync to TTS, so use a still plus audio on Fabric ($0.1875/s at 720p listed). Wan 3.0 is cheaper but will not match the voice.
- 10-minute avatar video from 60-second jobs: HeyGen cap vs Sume
HeyGen Creator renders up to 30 minutes. Sume avatar jobs run 4-60 seconds, so a 10-minute course is ten jobs joined by a $1.00 Timeline render.
- Five Sume image models are text-only. What to use with a photo
Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max take no reference images on Sume. Price and ratio of each, plus the cheapest edit-capable swap.
- Text or image start: Seedance 2.5 is $5.78 either way (10 s, 720p)
Starting a Sume video from an image does not change its price (model, resolution, ratio, seconds). Which models need an image, plus one H3 exception.
Written by Sume