ElevenLabs Avatars wants 3 to 5 reference images: what Sume needs
ElevenLabs recommends 3-5 reference images from different angles for Avatars. Sume Avatar 1.0 builds an avatar handle from one photo, a prompt, or simple props.
ElevenLabs recommends 3 to 5 reference images from different angles when you build an Avatar, and warns that a single image may be inconsistent. Sume Avatar 1.0 takes a different route: its avatar creation endpoint accepts one photo, a text prompt, or props for ethnicity, sex and age, and charges $0.95 flat per avatar.
What the vendor page says
The ElevenLabs documentation describes reference images or a text prompt as the inputs, with any library or cloned voice. It also notes that some Avatar models and reference-image uploads are restricted in the United States, so check availability for your region before you plan a workflow around uploads.
| Fact | Value |
|---|---|
| Recommended images | 3 to 5, different angles |
| Single image | May be inconsistent |
| Prompt input | Supported |
| Region note | Some models and uploads restricted in the US |
The Sume inputs
POST /v1/avatar-1.0/generate needs an avatar_handle and one input. With photo, you pass an image_url. With props you describe the person by ethnicity, sex and age. With prompt you describe them in words. The avatar creation docs list the fields.
Image URLs must be public HTTPS, and content types must match; localhost, private-network and signed URLs are rejected before the job is submitted.
Practical advice
Because the first frame is where a mismatch shows, preview before you pay for a long render. Use avatar video previews to see the opening frame, and override the final quality at generate-video without making a new preview.
If consistency across many clips matters more than a single image, test with the same handle across two or three scripts before you roll it out.
- Use a sharp, front-facing photo with even light.
- Keep one handle per recurring presenter.
- Re-run creation, not the video, when the face itself is wrong.
Sources
Related posts
More in Comparisons
- ElevenLabs Avatars has no API at launch: the Sume Avatar API
ElevenLabs says an Avatars API is planned but not live. If you need to automate avatar videos now, here is the Sume Avatar 1.0 route and what it needs.
- ElevenLabs Dubbing v2 cloning_strength 0-10: what Sume has instead
ElevenLabs Dubbing v2 has cloning_strength from 0 to 10, default 7. Sume has no such dial: you pick a voice per language and set speed. What that costs you.
- ElevenLabs maximum_text_length_per_request vs Sume max_characters
ElevenLabs now says to read maximum_text_length_per_request, not the old per-plan fields. Sume's TTS Router catalog publishes max_characters on every model row.
- ElevenLabs Music inpainting: redo one section vs a new Sume take
ElevenLabs Music v2 and v2.5 can regenerate one section of a song in the UI. Sume cannot edit a track: write a new brief for that part and generate again.
Written by Sume