Real-time AI avatar vs video avatar API: the difference

A real-time AI avatar talks live in a session; a video avatar API renders a finished file from a script. Which vendors sell which, and where Sume fits.

5 min readSume
All posts

A real-time AI avatar is a live face that listens and answers during a session, rendered while the conversation happens; a video avatar API renders a finished video file from a script you send, and you download it after the job completes. They are different products: HeyGen and Synthesia each sell a live avatar API next to their video API, while Sume's avatar API renders files only.

Vendor facts below come from HeyGen's and Synthesia's own docs, read on 2026-09-28. Sume facts come from Generate avatar video, Webhooks, and the Sume API reference.

How does a real-time avatar work?

Synthesia's docs put it plainly: an Interactive Avatar "isn't a video you generate and download. It's a live participant that Synthesia renders in real time" and places into a LiveKit room you control. Your own agent handles listening and turn-taking, and Synthesia renders the lip-synced face.

HeyGen's LiveAvatar works in two modes. In FULL mode HeyGen manages the whole conversation (speech-to-text, LLM, text-to-speech, turn-taking, memory). In LITE mode you keep your own agent and LiveAvatar supplies only the avatar layer.

Either way, both vendors bill live use by the minute, and the output is a live stream, not a file you keep.

Which vendors sell live avatars, and which render videos?

Listed alphabetically. Prices and limits are as each vendor states them, in the units it uses.

From HeyGen's Live Avatar, API pricing and LiveAvatar credits pages, Synthesia's API introduction and operational trust pages, and Sume's Generate avatar video, read 2026-09-28.
ProductKindWhat the vendor states
HeyGen LiveAvatarReal-timeFULL mode 2 credits per minute, LITE mode 1 credit per minute; plans and pricing are separate from API plans
HeyGen API videoRendered filePay-as-you-go API credits in USD, priced per minute and charged by actual seconds generated
Sume Avatar 1.0Rendered fileAsync job; one video covers an estimated 4-60 seconds at 720p
Synthesia Interactive Avatars APIReal-time"Synchronous and streaming, not a video file"; 10 credits per minute ($0.10/min); 100 concurrent sessions on paid plans
Synthesia Video APIRendered file"Asynchronous — you submit a job and poll or get a webhook when it's ready"

What does a video avatar API return?

A rendered avatar video is a job, not a stream. On Sume you send an avatar_handle and a script to POST /v1/avatar-1.0/talking-video, and in the default async mode the response carries the job's status_url and result_url right away. The sync mode waits at most 30 seconds and bounds the HTTP wait, not the job.

When the job completes, the result can include a public media.sume.com video. Webhooks are terminal events only (job.completed, job.failed, job.canceled); there are no progress or partial deliveries.

Each video covers an estimated 4-60 seconds of script, billed per second at $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). The request itself is covered in Talking avatar video API.

Can Sume run a live, talking-back avatar?

No. Sume's avatar video and speech routes are jobs you poll or get a webhook for; the API reference describes TTS 1.0 as an async job with poll or webhook, non-streaming. There is no session, room, or stream to join; What is an AI avatar? covers the same limit.

What Sume fits is a recorded answer: the words are known before anyone watches. FAQ videos with an AI avatar shows that pattern, one clip per question.

Which one do I need?

  • Pick a real-time avatar if a person talks to it and the answer depends on what they say: a support face on a website, a practice conversation, a kiosk. Budget for per-minute session time.
  • Pick a video avatar API if the script is written first: training steps, product explainers, FAQ answers, ads. You get a file you can review, caption, and host anywhere.
  • Check the live product's limits before you build. Synthesia's operational trust page lists no session transcript, recording, or webhook system, bust framing only, and no stock avatars in live sessions.
  • To have an AI agent decide what an avatar says, see AI avatar vs AI agent.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume