Custom AI avatar from a photo: allowlist or self-serve?

Google gates Live Avatar custom avatars behind enterprise allowlisting. Sume's avatar docs list a photo input on the create route. Request and rules.

4 min readSume
All posts

Google's Gemini 3.8 Live Avatar makes custom avatars from a high-quality reference image only through enterprise allowlisting. Sume's create-avatar docs list an Image input with no allowlist step described: you send input.type of photo and a public image_url to POST /v1/avatar-1.0/generate, poll the job, and use the returned handle.

This compares the access model, not quality. Google's page gives no price, and the Sume docs page does not promise a look-alike result from any given photo.

How does access differ?

Custom avatar from a reference image, read 2026-09-29.
Gemini Live AvatarSume avatar create
Reference imageHigh-quality reference imagephoto input with image_url
GateEnterprise allowlistingNone described on the docs page
Other ways inNot statedPrompt or Profile inputs
ResultA live, streamed avatarA reusable avatar handle for talking videos

What does the photo request look like?

The request needs a top-level avatar_handle and an input object. The handle may include a leading @; Sume stores it without one. Send an Idempotency-Key header as the docs example does.

curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-image-001" \
  -d '{
    "avatar_handle": "reference_presenter",
    "input": {
      "type": "photo",
      "image_url": "https://example.com/reference.png"
    }
  }'

What rules apply to the image URL?

  • image_url must be a fetchable public HTTPS image URL.
  • Localhost, private-network URLs, non-HTTPS URLs and non-image responses are rejected before generation is submitted.
  • Use a photo you have the right to use; the docs page does not cover consent, so that check is yours.

What if I would rather not use a photo?

Two other inputs exist on the same route. A Prompt input describes the avatar in text. A Profile input, sent as the props type, takes structured traits such as ethnicity, sex and age, which suits an app that already stores profile details for the avatar.

Either one avoids handling a person's photo at all, which can simplify consent questions. Each still creates a job and returns a handle you use in the same way.

Can one photo serve many videos?

Yes, in the sense the docs describe: the create job returns an avatar handle or resource id, and later avatar video requests reference it as avatar_handle. You create the avatar once and write a new script for each clip.

You can list your avatars with GET /v1/avatar-1.0/avatars, which is the read route the docs prefer. Nothing on the docs page describes editing the photo after creation, so make a new avatar if you need a different reference.

What happens after the job completes?

Each create request is a job. Poll GET /v1/jobs/{id}/status, then read /result. The returned avatar handle or resource id goes into the avatar video request as avatar_handle.

If you do not have a photo, the same route accepts a text Prompt or structured Profile traits instead.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume