Create an AI avatar from a reference image: URL rules and cost
Sume turns a public HTTPS photo into a reusable avatar for $0.95. The request, the URL checks that reject a bad image, and how to use the handle in videos.
To create an AI avatar from a reference image on Sume, send POST /v1/avatar-1.0/generate with an avatar_handle and an input of {"type": "photo", "image_url": "https://..."}. The image URL must be a fetchable public HTTPS image. Creation is a job, so you poll until it completes, then use the handle in POST /v1/avatar-1.0/talking-video. Each creation is a flat $0.95 per avatar on the current rate card.
Once you have the avatar, it is reusable: every later video costs the per-second video rate, not another creation fee.
The request
The handle is yours to name; a leading @ is allowed and is stored without it.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-image-001" \
-d '{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}'Why an image URL gets rejected
The docs say the URL must be a fetchable public HTTPS image, and that localhost, private-network, non-HTTPS and non-image responses are rejected before generation submission.
| Problem | What happens |
|---|---|
http:// instead of https:// | Rejected before generation |
localhost or a private-network address | Rejected before generation |
| URL returns a page, not an image (mismatched content type) | Rejected before generation |
| A link that needs a login or signature | Not stated in the docs; the image must be publicly fetchable, so test the link in a private browser window first |
Three ways to create, one handle to use
Whichever you choose, the result is an avatar_handle that you pass to the video route. List what you have with GET /v1/avatar-1.0/avatars, and read one with GET /v1/avatar-1.0/avatars/:id.
prompt: describe the avatar in text.props: structured traits such as ethnicity, sex and age, for apps that already store profile data.photo: a reference image, as above.
Choosing the photo
A reference photo works best when the face fills a good part of the frame, is lit evenly, looks toward the camera, and has nothing covering the mouth. Strong filters, heavy shadows or tiny faces in a wide shot give the model less to hold on to. If you only need a generic presenter and have no photo, the prompt and props types skip the image step and the URL checks entirely.
After creation
Poll GET /v1/jobs/:id/status until it is completed, then read GET /v1/jobs/:id/result. Then render a video. A 30-second plus clip with no product image is $7.35, so the first avatar plus one video is $8.30 in total (confirm with GET /v1/catalog).
Before you commit to a batch of videos, run the first one as a preview so you see the avatar framed in your chosen scene.
Limits and care
Use a photo you have the right to use. Many platforms restrict content that resembles a real person without consent, and the terms differ by destination, so check before you publish. Image quality drives avatar quality: a sharp, front-facing, evenly lit image is a better starting point than a cropped group photo. Avatar video output is 720p and each video is 4 to 60 seconds.
Sources
Related posts
More in Sume Avatar 1.0
- Face swap video_url rejected: signed and private URLs explained
Sume face swap needs a public HTTPS video_url. Signed or private URLs, localhost and provider task URLs are rejected before generation. How to host the clip.
- Lost an avatar video job id? List avatar videos instead
Sume has GET /v1/avatar-videos and GET /v1/avatar-videos/:id. How to find a render after a crash without resubmitting, and which status field to read.
- HeyGen Avatar 3.0 singing and 177 languages vs Sume Avatar 1.0
HeyGen Avatar 3.0 adds singing and 177+ languages. Sume Avatar 1.0 renders script-driven talking video, 4 to 60 seconds. What each one covers.
- Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)
Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.
Written by Sume