Create an AI avatar from a reference image: URL rules and cost

Sume turns a public HTTPS photo into a reusable avatar for $0.95. The request, the URL checks that reject a bad image, and how to use the handle in videos.

5 min readSume
All posts

To create an AI avatar from a reference image on Sume, send POST /v1/avatar-1.0/generate with an avatar_handle and an input of {"type": "photo", "image_url": "https://..."}. The image URL must be a fetchable public HTTPS image. Creation is a job, so you poll until it completes, then use the handle in POST /v1/avatar-1.0/talking-video. Each creation is a flat $0.95 per avatar on the current rate card.

Once you have the avatar, it is reusable: every later video costs the per-second video rate, not another creation fee.

The request

The handle is yours to name; a leading @ is allowed and is stored without it.

curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-image-001" \
  -d '{
    "avatar_handle": "reference_presenter",
    "input": {
      "type": "photo",
      "image_url": "https://example.com/reference.png"
    }
  }'

Why an image URL gets rejected

The docs say the URL must be a fetchable public HTTPS image, and that localhost, private-network, non-HTTPS and non-image responses are rejected before generation submission.

From Media inputs and Create new avatar, read 2026-10-01.
ProblemWhat happens
http:// instead of https://Rejected before generation
localhost or a private-network addressRejected before generation
URL returns a page, not an image (mismatched content type)Rejected before generation
A link that needs a login or signatureNot stated in the docs; the image must be publicly fetchable, so test the link in a private browser window first

Three ways to create, one handle to use

Whichever you choose, the result is an avatar_handle that you pass to the video route. List what you have with GET /v1/avatar-1.0/avatars, and read one with GET /v1/avatar-1.0/avatars/:id.

  • prompt: describe the avatar in text.
  • props: structured traits such as ethnicity, sex and age, for apps that already store profile data.
  • photo: a reference image, as above.

Choosing the photo

A reference photo works best when the face fills a good part of the frame, is lit evenly, looks toward the camera, and has nothing covering the mouth. Strong filters, heavy shadows or tiny faces in a wide shot give the model less to hold on to. If you only need a generic presenter and have no photo, the prompt and props types skip the image step and the URL checks entirely.

After creation

Poll GET /v1/jobs/:id/status until it is completed, then read GET /v1/jobs/:id/result. Then render a video. A 30-second plus clip with no product image is $7.35, so the first avatar plus one video is $8.30 in total (confirm with GET /v1/catalog).

Before you commit to a batch of videos, run the first one as a preview so you see the avatar framed in your chosen scene.

Limits and care

Use a photo you have the right to use. Many platforms restrict content that resembles a real person without consent, and the terms differ by destination, so check before you publish. Image quality drives avatar quality: a sharp, front-facing, evenly lit image is a better starting point than a cropped group photo. Avatar video output is 720p and each video is 4 to 60 seconds.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume