Create a Sume avatar from a photo: what image_url must satisfy
To make an avatar from a photo, send input type photo with a public HTTPS image_url. Sume rejects localhost, private, non-HTTPS and non-image URLs first.
Send input.type: "photo" with a fetchable public HTTPS image_url to POST /v1/avatar-1.0/generate. Before Sume submits the generation it rejects localhost, private-network and non-HTTPS URLs, and any response that is not an image. A signed or login-gated link will not work, so host the photo where an anonymous request can read it.
Three ways in, one route
Avatar creation uses a top-level avatar_handle and an input union. The handle may start with @, and Sume normalises it and stores it without the @. The union has three forms: prompt for text only, props for structured traits, and photo for a reference image.
Each request creates a job, not an avatar. Poll the job until it completes, then use the returned handle or resource id on the talking-video call.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-image-001" \
-d '{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}'What the URL check rejects
The same fetchability rule applies to every media field in the Avatar family, including product_image, scene.image_url and the background images in video_inputs. The shared guide is Media inputs.
| URL | Result |
|---|---|
http://... | Rejected, not HTTPS |
https://localhost/... or a private address | Rejected, not public |
| A link that returns HTML or a login page | Rejected, not an image |
| A public HTTPS image | Accepted |
Choosing photo, props or prompt
Whichever you choose, keep the handle stable. Every later call refers to the avatar by avatar_handle, and a reserved prefix is refused, so pick a name that belongs to your own series.
- Use
photowhen a real reference image exists and you want the avatar to follow it. - Use
propswhen your app already stores traits and you want repeatable, structured creation. - Use
promptwhen you only have a description.
Before you upload a face
A reference photo of a real person is a decision about consent, not just about format. Confirm that the person agreed to the use before you send the image, and keep the record with the handle. Sume validates the URL; it cannot validate the permission.
After the avatar is ready, read it with GET /v1/avatar-1.0/avatars/{id}, then make a preview of the first video before the full render.
Older paths
POST /v1/models/sume/avatar-1.0/generate/runs and the legacy POST /v1/models/sume/avatar/v1.0/runs accept the same body. For new integrations the docs prefer /v1/avatar-1.0/generate.
Related posts
More in Sume Avatar 1.0
- Does an AI avatar presenter make a Short original?
An avatar is a delivery method; YouTube's pages ask for original substance. How to give an Avatar 1.0 Short your own angle within its 60-second job limit.
- Face swap beta: motion stage runs on Kling, source audio muxed back
Sume's Avatar Face Swap beta runs its motion stage on the Kling 3.0 motion control queue, strips the source audio, then muxes the original audio back in.
- Full-duplex avatar listens while it speaks: script a clip instead
A full-duplex avatar hears you mid-sentence. If your content is scripted, build similar beats into one Sume avatar clip with silence scenes. Curl included.
- Full-duplex or script-driven avatar: six questions before you pick
Tavus Griffin-Lite is a full-duplex video conversation model in research preview. Six questions that separate it from Sume Avatar 1.0 talking video.
Written by Sume