Atmee LiveKit avatar from one portrait vs a Sume photo avatar clip
Atmee turns one portrait into a live LiveKit avatar. Sume turns one photo into a reusable avatar handle and finished clips. Pick by who waits on the face.
Atmee's LiveKit plugin builds a talking-head avatar from a single portrait (jpg, png or webp) and renders it live in a LiveKit room. Sume also starts from one portrait, but the result is a reusable avatar handle and a finished MP4, not a participant in a call. Choose by whether a person is waiting on the face in real time.
How the Atmee plugin starts
According to the plugin's README, creating an avatar needs a name and one portrait, given as a file path, raw bytes or an HTTPS URL. The avatar is render-ready immediately. You install livekit-plugins-atmee (Python 3.10 or later, livekit-agents 1.6.8 or later), set an ATMEE_API_KEY, and the avatar joins the room as its own participant, lip-synced to your agent's speech.
The README says start() returns in about a second once the render worker acknowledges, and the avatar's tracks appear a few seconds later. The licence is Apache-2.0.
How Sume starts
Sume's Avatar 1.0 has two steps. First, POST /v1/avatar-1.0/generate creates an avatar from a prompt, a profile or a photo (input.type: "photo" with a public HTTPS image_url). The job completes and the avatar becomes a reusable resource in your workspace under a handle. Second, POST /v1/avatar-1.0/talking-video takes that avatar_handle and a script and returns a talking video.
Creating an avatar costs a flat $0.95. A talking video is priced per second by quality: standard $0.184, plus (default) $0.245, max $0.55 per second without a product image. Scripts must estimate to 4 to 60 seconds.
| Question | Atmee plugin | Sume Avatar 1.0 |
|---|---|---|
| Input | One portrait, jpg/png/webp | One photo for avatar create, then a script |
| Result | Live participant in a LiveKit room | Reusable handle, then an MP4 |
| Ready time | Render-ready immediately, tracks in a few seconds | Avatar create is a job; poll until completed |
| Cost shape | Per minute while the avatar is in the room | $0.95 once, then per second of video |
| 60 s of output | Minutes of session time | 60 x $0.245 = $14.70 at plus quality |
Which to pick
If your product is a voice agent that should have a face during a call, the plugin is built for that. If your product is the same presenter appearing in many scheduled videos, a handle is the better fit, since every video reuses the same face. The cost of a brand-consistent presenter walks through twelve videos on one avatar.
Keep the source portrait sharp and front-facing in either case. Both systems can only animate what the photo shows.
Sources
Related posts
More in Comparisons
- Best API for text to speech: six checks to run before you pick
Compare TTS APIs on price per million characters, request cap, latency, languages, controls and word timing, with MAI-Voice-2.1 and Sume TTS 1.0 numbers.
- Beyond Presence 50 credits a minute vs Sume lip sync per second
Beyond Presence bills live speech-to-video at 50 credits a minute. Sume bills a rendered lip-sync clip per second. Here is how to compare the two units.
- bitHuman local .imx or cloud avatar vs a hosted Sume avatar job
bitHuman renders avatars locally from .imx files or in its cloud, which changes your CPU bill. A hosted Sume avatar job moves the render off your agent.
- Can an AI avatar see me on a video call? Griffin perception vs Sume
Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.
Written by Sume