Gemini Omni avatar with your own voice vs Sume Avatar 1.0 video
Google says Gemini Omni can make an avatar video in your own voice. Sume Avatar 1.0 takes a script, 4 to 60 seconds, 720p. What differs.
Google says Gemini Omni lets you make videos with a digital avatar of yourself using your own voice, in the Gemini app. Sume's Avatar 1.0 makes talking videos from a ready avatar and a script in 4 to 60 seconds at 720p, with a voice that follows the script text; the docs describe no way to supply your own recorded voice.
Google's claim is from Introducing Gemini Omni and the Omni Flash API page; Sume's contract is the avatar video docs. All were read 2026-10-02.
What does Google say Omni does with voice?
The Gemini Omni announcement says you can create videos with your own voice using digital avatars, and notes that editing speech and audio is something Google is still testing before offering it widely. On the developer side, the Omni Flash API page says audio references are unsupported in the current version, that voice editing is not supported, and that you cannot extend an uploaded video in which someone is talking to add new dialogue.
So the avatar-with-your-voice feature described in the announcement is a product feature in Google's apps, while the API page lists no audio-reference input. I found no API field for it in the pages I fetched.
What does Sume's Avatar 1.0 take?
A request to POST /v1/avatar-1.0/talking-video references a ready avatar by avatar_handle and provides exactly one of script or video_inputs. Sume estimates the target duration and accepts 4 to 60 seconds. quality is standard, plus (the default) or max; aspect_ratio is one of 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 by default; resolution is currently 720p.
Multi-scene plans use ordered video_inputs, where each scene has a voice of type text with a script or input_text, or a silence beat with a required duration. Current execution supports one resolved avatar per final video. The docs do not describe uploading a voice sample or recorded audio for the avatar to use.
| Question | Gemini Omni | Sume Avatar 1.0 |
|---|---|---|
| Own-voice avatar | Announced for Gemini app | Not in the docs |
| Audio reference input | Unsupported on the API page | No reference audio field |
| Length | 3 to 10 s per clip (API) | 4 to 60 s per video |
| Resolution | 360p to 4K | 720p |
| Voice source | Not specified for API | Text script in the request |
Which should you use?
If your goal is a clip of yourself speaking in your own voice and you are a Google AI subscriber, the Gemini app is where Google says the feature lives. If your goal is a repeatable spokesperson video from a script, such as product explainers made per SKU, Avatar 1.0's one-call script-to-video route with an idempotency key is built for that. For videos longer than Omni's 10-second clips, 4 to 60 seconds in one job is the practical difference.
Sume does not clone a user's voice in the Avatar route, and I have not claimed otherwise. Check the avatar docs for changes before you depend on that.
What does a request look like?
A single-script talking video with a ready avatar. Replace the handle with one of yours; the call returns a job to poll.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-video-001" \
-d '{
"avatar_handle": "your_avatar_handle",
"script": "Meet the new travel mug: it keeps drinks hot for twelve hours.",
"aspect_ratio": "9:16",
"quality": "plus"
}'How do the two handle dialogue in long videos?
Google's page says multi-turn voice extension works for generated videos, and that new dialogue cannot be added by extending an uploaded video in which someone is already talking. Clips are 3 to 10 seconds, with a 40-second ceiling when extended.
Sume's route plans the full script up front. A 45-second script is one job inside the 4 to 60 second window, not a chain of extensions, and a script longer than that is split into several jobs. Silence beats let you leave room for a cutaway shot you add later in a timeline.
What is still unknown?
Google's announcement does not state which countries, plans or length limits apply to the own-voice avatar, and I found no API for it. Revisit the Omni page before you plan a build around either claim.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen Avatar 3.0 singing and 177 languages vs Sume Avatar 1.0
HeyGen Avatar 3.0 adds singing and 177+ languages. Sume Avatar 1.0 renders script-driven talking video, 4 to 60 seconds. What each one covers.
- Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)
Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.
- Regenerate avatar preview stills, or start a new preview?
Regenerate refreshes first-frame stills from the stored preview request. Changing script, avatar, scene or aspect ratio needs a new preview. The full rule.
- One avatar handle, three platform cuts: keep the character consistent
Keep one presenter across LinkedIn, Snapchat and Pinterest by reusing a single avatar_handle and rendering each cut at the right ratio and length with Sume.
Written by Sume