Gemini 3.8 Live Avatar is enterprise-only; what Sume offers for clips
Google's Live Avatar in Gemini 3.8 Live is a real-time persona for Gemini Enterprise in 97 languages. Sume's Avatar 1.0 renders English clips. How they differ.
Google's Live Avatar, announced September 24, 2026 with Gemini 3.8 Live, is a real-time spoken persona that you reach through Gemini Enterprise, and Sume has no live or real-time avatar to match it. Sume's Avatar 1.0 renders a recorded talking video from a script in English only, in 4 to 60 seconds, as a job you poll.
Google's blog post, read 2026-10-11, describes the feature as pairing real-time video generation with speech for an assistant with "a dynamic visual persona" and "precise lip-syncing, natural expressions, and fluid turn-taking". The Sume side comes from Generate avatar video.
What Google says about access
The post says the feature is currently accessible through Gemini Enterprise only, and available through an API documented in the Google Cloud console. It lists 97 languages with speech-to-speech handling, preset avatars, and custom avatars from reference images, where custom creation is limited to enterprise allowlisting for now. Output carries SynthID watermarking in both audio and video. The post gives no pricing.
That means there is nothing to quote about cost, and this post does not try to.
One detail worth noting is the allowlisting step. Custom avatars from reference images are not self-serve in the Google post; creating one currently requires enterprise allowlisting. By contrast, a Sume avatar is created with a single API call from a text prompt, structured traits or a photo, for a flat $0.95, and the avatar can be used as soon as its job completes.
Live persona versus rendered clip
The products are built for different jobs, so the table lists what each documents rather than naming a winner.
| Question | Gemini 3.8 Live Avatar | Sume Avatar 1.0 |
|---|---|---|
| Interaction | Real-time, turn-taking conversation | Script in, finished video out |
| Access | Gemini Enterprise; custom avatars by allowlist | API key; avatar created from a prompt, traits or photo |
| Languages | 97 per Google's post | English only |
| Length | Not stated in the post | 4-60 seconds per job |
| Price stated | None in the post | Published per second by tier |
| Provenance mark | SynthID watermark | Not described in the avatar docs I read |
When a recorded clip is the right tool
If the viewer will never talk back, you do not need a live model. A product-launch message, an onboarding step, a reminder or a social ad is a fixed script. A recorded clip lets a human read the words before anyone sees them, lets you check the first frame in advance with a preview, and gives a stable file with a known price. A live persona is the right pick for an assistant that must answer unscripted questions.
Because Sume's avatar speaks English only, a clip for Spanish or Japanese viewers is not something Avatar 1.0 can make; the English-only note in the Spanish question post lists what to do instead.
Think about review, too. In a live conversation nobody can approve what the persona says before the viewer hears it, so the controls sit in the model and its guardrails. In a recorded clip the control is editorial: you read the script, approve the first frame, and check the transcript of the finished video before publishing. Teams with a compliance step usually prefer the second model for anything regulated, and the first for open-ended help.
What to do this week
If you are on Gemini Enterprise, ask your admin whether Live Avatar is enabled for your organization and what the allowlisting process is for custom avatars. If you need a clip now in English, create an avatar once, write the script, approve the first frame and render. Keep your live-agent evaluation and your clip production as separate workstreams; they will share almost no code.
Whatever you publish, label the synthetic presenter. Google marks its output itself; with a Sume clip you add the disclosure line to the caption or description.
Sources
Related posts
More in Comparisons
- Merchant Center Product Studio video vs the Sume video API
Product Studio saves 720p, caps images at 32 MB and works in six countries. When a Sume video API request fits better, read from Google's page and Sume docs.
- HeyGen enterprise writes rise to 30/s; how Sume limits submits
HeyGen raised enterprise limits: writes 30 per second, polling 1,500 per minute. Sume publishes no per-second number; it uses plan queues and rate headers.
- HeyGen instant voice clone audio in Sume? Only Sume-hosted audio
HeyGen's new instant voice clone can output TTS audio, but Sume's talking-still routes take only Sume-hosted audio up to 10 MB. What it blocks and what works.
- Is AI video 'Full HD' native or upscaled? Kandinsky 6.0 vs Sume rows
Kandinsky 6.0 renders 864x480 and adds Full HD with a plug-in. Sume's H3 Max 1080p is a latent refinement from 768p. What that means for price and detail.
Written by Sume