Gemini Live Avatar vs a video avatar API: live or rendered?

Gemini 3.8 Live Avatar streams a lip-synced face in a conversation. Sume's avatar video renders a finished clip from a script. Which job needs which.

4 min readSume
All posts

Gemini Live Avatar and a video avatar API solve different problems. Google describes Live Avatar as lip-synced streaming video on a live dialogue model, so the face answers while a person talks. Sume's avatar video is the other shape: you send a script, poll a job, and get back a finished clip you can review before anyone sees it.

Google's announcement is dated 2026-09-24 and says the feature is available in Gemini Enterprise. The page snapshot read for this post lists no price, so this post compares shapes, not cost.

What does each one give you?

Live Avatar and Sume avatar video compared, read 2026-09-29.
Gemini Live AvatarSume avatar video
OutputLip-synced streaming video on a live dialogue modelA script-driven talking video from a ready avatar
What decides the wordsThe live conversationA script or ordered video_inputs you send
LengthGoogle's page gives no clip lengthEstimated 4 to 60 seconds per job
DeliveryStreamedJob you poll, or a webhook

How does Sume's avatar video work?

The avatar video docs describe it as turning a ready avatar into a script-driven talking video. You call POST /v1/avatar-1.0/talking-video with an avatar_handle and exactly one of script or video_inputs. The response is a job; you read /v1/jobs/{id}/status and then /result.

Because the words are fixed before rendering, you can preview first-frame stills, add captions, and reject a take before publishing. A live stream has no such review step by nature.

Can Sume answer a viewer in real time?

No. Nothing in the avatar video docs describes a session or a stream; the route creates a job. If a viewer's question decides the answer, a live avatar is the right shape. If you know the question in advance, render one clip per answer and serve the file.

What does a rendered avatar let you control?

The request takes quality of standard, plus (the default) or max, an aspect_ratio of 1:1, 3:4, 9:16 (the default), 4:3 or 16:9, and a resolution that is currently 720p. Optional scene fields set the background, either from a prompt or a public HTTPS photo.

That control is the trade for giving up live turns: each choice is made before the render starts, and you can review first-frame stills before paying for the full render.

What is still unknown about Live Avatar?

Google's announcement, as read on 2026-09-29, names the availability (Gemini Enterprise), the streaming shape and the language and watermark facts covered in Gemini Live Avatar: 97 languages and SynthID. It gives no price and no clip length, so any cost comparison would be invented. Read the vendor page again before you plan around it.

Which one should I pick?

  • Pick a live avatar when the reply depends on what a person says next.
  • Pick rendered clips when the script exists first: product explainers, FAQ answers, onboarding steps.
  • Check Google's page for who can enable Live Avatar; it lists Gemini Enterprise.

Sources

Related posts

More in Models

All Models posts

Written by Sume