Gemini app video avatar: the @username prompt vs a Sume avatar handle
In the Gemini app you add your avatar to a video by typing @ and your Google username, 18+ only. Sume uses an avatar_handle and a photo input. What differs.
In the Gemini app you put your avatar in a video by typing @ followed by your Google username in the prompt, and you must be 18 or older. On Sume you create an avatar once with an avatar_handle and then reference that handle in a talking-video request of 4 to 60 seconds.
Google's side is from the Gemini Apps Help page and the Gemini Omni overview, both read on 2026-10-02. Sume's side is from its Create new avatar and Generate avatar video docs.
How does the Gemini app avatar work?
The help page says you include your avatar by typing the at sign and your Google username in the prompt, and that audio can be generated with the video. The overview describes an avatar as optional, 18 and over, and says it lets you make videos that look and sound like you. Requirements for video in the app also include a Google AI plan on personal accounts and a qualifying Workspace license on work or school accounts.
Neither page gives a clip length for avatar videos, so treat the 10-second figure on the overview as the general video cap, not as an avatar-specific rule.
How does a Sume avatar work?
Sume has three ways to create an avatar: a text prompt, structured traits, or a reference image. The image path takes input.type set to photo and a public HTTPS image_url. The handle may include a leading at sign, and Sume stores it without one. Creation is a job, so you poll it, then use the handle.
A talking video then takes avatar_handle plus exactly one of script or video_inputs. Sume estimates the duration and accepts 4 to 60 seconds. aspect_ratio defaults to 9:16, and resolution is currently 720p. quality is standard, plus or max, with plus as the default.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-image-001" \
-d '{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}'What are the real differences?
The age row deserves care. Sume's docs do not state an age rule for avatars, and that is not the same as having none. If you build a product around people's faces, put consent and age checks in your own flow.
| Question | Gemini app | Sume Avatar 1.0 |
|---|---|---|
| How you refer to the avatar | @ plus your Google username in the prompt | avatar_handle in the request |
| Who it is made from | You, as an optional feature | A prompt, traits or a reference photo |
| Age rule | 18 or older | Not stated in the avatar docs I read |
| Clip length | Not stated for avatars | 4 to 60 seconds |
| Resolution | Not stated for avatars | 720p |
| Voice | Audio generated with the video | A script, or per-scene voice and silence in video_inputs |
Which one fits which job?
Use the app when the video is about you, you are an adult on a plan that includes video, and a 10-second clip is enough. Use Sume when you need a reusable character with a handle, a script longer than 10 seconds, or many videos in a batch. Sume does not clone your Google account's avatar, and there is no link between the two.
What would a repeatable avatar workflow look like?
Create the avatar once, then spend your time on scripts. Write each script to fit the 4 to 60 second window, because Sume estimates the duration from the script and rejects longer ones; split a long piece into several jobs. For scene control use video_inputs, where each scene has an id, a voice of text or silence with a duration, and a background.
Check each result before it goes out. A talking video built from a reference photo is still a synthetic person, and you should check the disclosure rules of every platform you publish on. Keep the job id next to each published file so you can prove how it was made.
Keep your own file of consent. If the face is yours, say so in the record. If it is someone else's, get written permission first. Neither Google's pages nor Sume's docs do that for you.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen Avatar 3.0 singing and 177 languages vs Sume Avatar 1.0
HeyGen Avatar 3.0 adds singing and 177+ languages. Sume Avatar 1.0 renders script-driven talking video, 4 to 60 seconds. What each one covers.
- Korean avatar video captions: fix caption_hangul_text_latin_style
A Korean script with captions style slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style. Use a Hangul style instead.
- Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)
Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.
- Regenerate avatar preview stills, or start a new preview?
Regenerate refreshes first-frame stills from the stored preview request. Changing script, avatar, scene or aspect ratio needs a new preview. The full rule.
Written by Sume