ElevenLabs Avatars wants 3 to 5 reference images: what Sume needs

ElevenLabs recommends 3-5 reference images from different angles for Avatars. Sume Avatar 1.0 builds an avatar handle from one photo, a prompt, or simple props.

4 min readSume
All posts

ElevenLabs recommends 3 to 5 reference images from different angles when you build an Avatar, and warns that a single image may be inconsistent. Sume Avatar 1.0 takes a different route: its avatar creation endpoint accepts one photo, a text prompt, or props for ethnicity, sex and age, and charges $0.95 flat per avatar.

What the vendor page says

The ElevenLabs documentation describes reference images or a text prompt as the inputs, with any library or cloned voice. It also notes that some Avatar models and reference-image uploads are restricted in the United States, so check availability for your region before you plan a workflow around uploads.

ElevenLabs reference-image guidance (read 2026-10-02)
FactValue
Recommended images3 to 5, different angles
Single imageMay be inconsistent
Prompt inputSupported
Region noteSome models and uploads restricted in the US

The Sume inputs

POST /v1/avatar-1.0/generate needs an avatar_handle and one input. With photo, you pass an image_url. With props you describe the person by ethnicity, sex and age. With prompt you describe them in words. The avatar creation docs list the fields.

Image URLs must be public HTTPS, and content types must match; localhost, private-network and signed URLs are rejected before the job is submitted.

Practical advice

Because the first frame is where a mismatch shows, preview before you pay for a long render. Use avatar video previews to see the opening frame, and override the final quality at generate-video without making a new preview.

If consistency across many clips matters more than a single image, test with the same handle across two or three scripts before you roll it out.

  • Use a sharp, front-facing photo with even light.
  • Keep one handle per recurring presenter.
  • Re-run creation, not the video, when the face itself is wrong.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume