ElevenLabs Avatars has no API at launch: the Sume Avatar API
ElevenLabs says an Avatars API is planned but not live. If you need to automate avatar videos now, here is the Sume Avatar 1.0 route and what it needs.
ElevenLabs Avatars does not offer an API at launch. The ElevenLabs documentation says an API for Avatars is planned, and the product is used through its interface on paid plans. If you must generate avatar videos from code or an agent today, Sume Avatar 1.0 exposes them as plain HTTP: POST /v1/avatar-1.0/talking-video returns a job you can poll or receive by webhook.
What ElevenLabs documents
Avatars is available on all paid plans, takes reference images or a text prompt, and can use any voice from the library, including cloned voices. The lip-sync model is selected automatically. Credits follow the Image & Video pricing and are shared across ElevenCreative tools.
| Fact | Value |
|---|---|
| Plans | All paid plans |
| API | Planned, not at launch |
| Inputs | Reference images or a text prompt |
| Voice | Any library voice, including clones |
| Lip-sync model | Auto-selected |
The Sume route
An avatar video needs an avatar handle plus exactly one of script or video_inputs. The default communication mode is async and returns 202. Poll /v1/jobs/:id/status or /result, or subscribe a webhook that is signed with HMAC SHA256.
Send an idempotency key with the request so a retry cannot create a second paid job.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: launch-clip-001" \
-d '{"avatar_handle":"product_host","quality":"standard","script":"Welcome. Here is the new feature in thirty seconds."}'
When the interface is enough
If a person is making a handful of clips and wants the ElevenLabs voice library, the interface may be all you need. The API matters when clips are made by a scheduler, a CMS or an agent, and you want admission, spend and retries to be explicit.
Re-check the ElevenLabs page before you plan around any future API; the vendor has not given a date.
Sources
Related posts
More in Comparisons
- ElevenLabs Dubbing v2 cloning_strength 0-10: what Sume has instead
ElevenLabs Dubbing v2 has cloning_strength from 0 to 10, default 7. Sume has no such dial: you pick a voice per language and set speed. What that costs you.
- ElevenLabs maximum_text_length_per_request vs Sume max_characters
ElevenLabs now says to read maximum_text_length_per_request, not the old per-plan fields. Sume's TTS Router catalog publishes max_characters on every model row.
- ElevenLabs Music inpainting: redo one section vs a new Sume take
ElevenLabs Music v2 and v2.5 can regenerate one section of a song in the UI. Sume cannot edit a track: write a new brief for that part and generate again.
- ElevenLabs Music length: 3 seconds to 5 or 10 minutes vs Sume
ElevenLabs' two pages disagree on max Music length: 5 minutes on the capabilities page, 600000 ms on the Compose reference. Sume has no length field at all.
Written by Sume