Real-time AI avatar vs video avatar API: the difference
A real-time AI avatar talks live in a session; a video avatar API renders a finished file from a script. Which vendors sell which, and where Sume fits.
A real-time AI avatar is a live face that listens and answers during a session, rendered while the conversation happens; a video avatar API renders a finished video file from a script you send, and you download it after the job completes. They are different products: HeyGen and Synthesia each sell a live avatar API next to their video API, while Sume's avatar API renders files only.
Vendor facts below come from HeyGen's and Synthesia's own docs, read on 2026-09-28. Sume facts come from Generate avatar video, Webhooks, and the Sume API reference.
How does a real-time avatar work?
Synthesia's docs put it plainly: an Interactive Avatar "isn't a video you generate and download. It's a live participant that Synthesia renders in real time" and places into a LiveKit room you control. Your own agent handles listening and turn-taking, and Synthesia renders the lip-synced face.
HeyGen's LiveAvatar works in two modes. In FULL mode HeyGen manages the whole conversation (speech-to-text, LLM, text-to-speech, turn-taking, memory). In LITE mode you keep your own agent and LiveAvatar supplies only the avatar layer.
Either way, both vendors bill live use by the minute, and the output is a live stream, not a file you keep.
Which vendors sell live avatars, and which render videos?
Listed alphabetically. Prices and limits are as each vendor states them, in the units it uses.
| Product | Kind | What the vendor states |
|---|---|---|
| HeyGen LiveAvatar | Real-time | FULL mode 2 credits per minute, LITE mode 1 credit per minute; plans and pricing are separate from API plans |
| HeyGen API video | Rendered file | Pay-as-you-go API credits in USD, priced per minute and charged by actual seconds generated |
| Sume Avatar 1.0 | Rendered file | Async job; one video covers an estimated 4-60 seconds at 720p |
| Synthesia Interactive Avatars API | Real-time | "Synchronous and streaming, not a video file"; 10 credits per minute ($0.10/min); 100 concurrent sessions on paid plans |
| Synthesia Video API | Rendered file | "Asynchronous — you submit a job and poll or get a webhook when it's ready" |
What does a video avatar API return?
A rendered avatar video is a job, not a stream. On Sume you send an avatar_handle and a script to POST /v1/avatar-1.0/talking-video, and in the default async mode the response carries the job's status_url and result_url right away. The sync mode waits at most 30 seconds and bounds the HTTP wait, not the job.
When the job completes, the result can include a public media.sume.com video. Webhooks are terminal events only (job.completed, job.failed, job.canceled); there are no progress or partial deliveries.
Each video covers an estimated 4-60 seconds of script, billed per second at $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). The request itself is covered in Talking avatar video API.
Can Sume run a live, talking-back avatar?
No. Sume's avatar video and speech routes are jobs you poll or get a webhook for; the API reference describes TTS 1.0 as an async job with poll or webhook, non-streaming. There is no session, room, or stream to join; What is an AI avatar? covers the same limit.
What Sume fits is a recorded answer: the words are known before anyone watches. FAQ videos with an AI avatar shows that pattern, one clip per question.
Which one do I need?
- Pick a real-time avatar if a person talks to it and the answer depends on what they say: a support face on a website, a practice conversation, a kiosk. Budget for per-minute session time.
- Pick a video avatar API if the script is written first: training steps, product explainers, FAQ answers, ads. You get a file you can review, caption, and host anywhere.
- Check the live product's limits before you build. Synthesia's operational trust page lists no session transcript, recording, or webhook system, bust framing only, and no stock avatars in live sessions.
- To have an AI agent decide what an avatar says, see AI avatar vs AI agent.
Sources
- Generate avatar video
- Webhooks
- Sume API reference
- API reference
- HeyGen Live Avatar (read 2026-09-28)
- HeyGen API pricing explained (read 2026-09-28)
- LiveAvatar credits and subscriptions (read 2026-09-28)
- Synthesia APIs introduction (read 2026-09-28)
- Synthesia Interactive Avatar concepts (read 2026-09-28)
- Synthesia Interactive Avatars operational trust (read 2026-09-28)
Related posts
More in Sume Avatar 1.0
- Runway Characters API: avatars, live sessions and videos
Runway Characters builds avatars from one image, a voice and a personality, for live sessions or rendered videos. Endpoints, limits and pricing.
- Synthesia alternatives with an API: access, pricing, limits
Argil, Creatify, HeyGen and Sume each document an avatar video API. They differ in how API access is sold, the price unit and the length limits.
- Synthesia personal avatar vs studio avatar: what differs
A Synthesia personal avatar is self-serve, from one photo, ready in minutes; a studio avatar is a green-screen shoot sold as a paid add-on.
- Talking head vs B-roll: what's the difference?
A talking head is the shot of someone speaking to camera; B-roll is the footage you cut to while their voice keeps playing. How the two fit together.
Written by Sume