Gemini Omni avatar with your own voice vs Sume Avatar 1.0 video

Google says Gemini Omni can make an avatar video in your own voice. Sume Avatar 1.0 takes a script, 4 to 60 seconds, 720p. What differs.

5 min readSume
All posts

Google says Gemini Omni lets you make videos with a digital avatar of yourself using your own voice, in the Gemini app. Sume's Avatar 1.0 makes talking videos from a ready avatar and a script in 4 to 60 seconds at 720p, with a voice that follows the script text; the docs describe no way to supply your own recorded voice.

Google's claim is from Introducing Gemini Omni and the Omni Flash API page; Sume's contract is the avatar video docs. All were read 2026-10-02.

What does Google say Omni does with voice?

The Gemini Omni announcement says you can create videos with your own voice using digital avatars, and notes that editing speech and audio is something Google is still testing before offering it widely. On the developer side, the Omni Flash API page says audio references are unsupported in the current version, that voice editing is not supported, and that you cannot extend an uploaded video in which someone is talking to add new dialogue.

So the avatar-with-your-voice feature described in the announcement is a product feature in Google's apps, while the API page lists no audio-reference input. I found no API field for it in the pages I fetched.

What does Sume's Avatar 1.0 take?

A request to POST /v1/avatar-1.0/talking-video references a ready avatar by avatar_handle and provides exactly one of script or video_inputs. Sume estimates the target duration and accepts 4 to 60 seconds. quality is standard, plus (the default) or max; aspect_ratio is one of 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 by default; resolution is currently 720p.

Multi-scene plans use ordered video_inputs, where each scene has a voice of type text with a script or input_text, or a silence beat with a required duration. Current execution supports one resolved avatar per final video. The docs do not describe uploading a voice sample or recorded audio for the avatar to use.

Avatar-with-voice features compared, read 2026-10-02
QuestionGemini OmniSume Avatar 1.0
Own-voice avatarAnnounced for Gemini appNot in the docs
Audio reference inputUnsupported on the API pageNo reference audio field
Length3 to 10 s per clip (API)4 to 60 s per video
Resolution360p to 4K720p
Voice sourceNot specified for APIText script in the request

Which should you use?

If your goal is a clip of yourself speaking in your own voice and you are a Google AI subscriber, the Gemini app is where Google says the feature lives. If your goal is a repeatable spokesperson video from a script, such as product explainers made per SKU, Avatar 1.0's one-call script-to-video route with an idempotency key is built for that. For videos longer than Omni's 10-second clips, 4 to 60 seconds in one job is the practical difference.

Sume does not clone a user's voice in the Avatar route, and I have not claimed otherwise. Check the avatar docs for changes before you depend on that.

What does a request look like?

A single-script talking video with a ready avatar. Replace the handle with one of yours; the call returns a job to poll.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-video-001" \
  -d '{
    "avatar_handle": "your_avatar_handle",
    "script": "Meet the new travel mug: it keeps drinks hot for twelve hours.",
    "aspect_ratio": "9:16",
    "quality": "plus"
  }'

How do the two handle dialogue in long videos?

Google's page says multi-turn voice extension works for generated videos, and that new dialogue cannot be added by extending an uploaded video in which someone is already talking. Clips are 3 to 10 seconds, with a 40-second ceiling when extended.

Sume's route plans the full script up front. A 45-second script is one job inside the 4 to 60 second window, not a chain of extensions, and a script longer than that is split into several jobs. Silence beats let you leave room for a cutaway shot you add later in a timeline.

What is still unknown?

Google's announcement does not state which countries, plans or length limits apply to the own-voice avatar, and I found no API for it. Revisit the Omni page before you plan a build around either claim.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume