Make an AI avatar from a selfie: the photo URL rules on Sume
Sume builds an avatar from a prompt, a profile or a photo. The photo must be a public HTTPS image URL; localhost, private and non-image URLs are rejected.
To make an avatar from a selfie on Sume, send the photo as input.image_url with type photo. The URL must be a fetchable public HTTPS image. Sume rejects localhost, private-network, non-HTTPS and non-image responses before it submits the generation, so a local file path or a signed private link will not work.
Three ways to create an avatar
The Create new avatar page (read 2026-10-08) lists three inputs, all sent to POST /v1/avatar-1.0/generate with a top-level avatar_handle.
| Method | API input type | What you provide |
|---|---|---|
| Prompt | prompt | A text description |
| Profile | props | Structured traits such as ethnicity, sex and age |
| Image | photo | A public HTTPS image_url |
Photo requirements in practice
The checks happen before submission, so failures are quick and cheap. Walk through this list before you call the API:
- The scheme is https, not http.
- The host is public: not localhost and not a private network address.
- The response is an image, not an HTML page that wraps an image.
- The link works without a login or an expiring signature.
- You have the right to use the face in the photo.
After the request
Avatar creation is job based. Sume returns a job, you poll /v1/jobs/{id}/status until it completes, then fetch /result. The finished avatar becomes a reusable resource in your workspace, listed at GET /v1/avatar-1.0/avatars. Reuse the handle in every later talking-video call.
Handles may start with @; Sume normalizes the handle and stores it without the @. Pick something stable, such as reference_presenter, so your app does not depend only on a generated id.
If you only have a photo on a private drive, upload it to a public host first. Media field rules are collected on Media inputs. Do not submit a photo of a person without their consent; the face will appear in published videos.
A request body
The docs example uses this shape for an image-based avatar.
{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}Choosing between the three inputs
Use a prompt when you want a generic presenter and have no source image. Use a profile (props) when your app already stores structured traits and needs repeatable avatars. Use a photo when the avatar should look like a specific reference. The photo route is the one with the most checks, because the image must pass the URL rules before the job is accepted.
Plan for consent and rights. A face in a photo is personal data, and a synthetic version of it can be published to many people. Keep written permission with your records, and avoid using images of people who have not agreed.
If the creation job fails, read the job error and try a different image rather than resubmitting the same URL. Use an Idempotency-Key for each paid creation attempt, as the Generation admission page recommends.
Sources
Related posts
More in Sume Avatar 1.0
- One Sume avatar handle across talking video, TTS, stills and lip sync
A Sume avatar handle is accepted by talking-video, TTS voice selection, face-in-still images and the lip-sync routes. What each uses from it and what it costs.
- Preview an AI avatar video before paying for the full render
Sume avatar video previews make first-frame stills only. Approve them, then call generate-video; you can change the final quality tier without a new preview.
- Put the AI disclosure in scene one: Sume avatar video_inputs recipe
A copyable Sume Avatar 1.0 request with a 3-second disclosure scene and a 12-second message. The 15 seconds cost $2.76 at standard, $3.68 at plus.
- Video call avatar or scripted talking video: which do you need?
A real-time avatar answers people live; a scripted talking video is a file you render and review first. How to choose, and what Sume Avatar 1.0 covers.
Written by Sume