Atmee LiveKit avatar from one portrait vs a Sume photo avatar clip

Atmee turns one portrait into a live LiveKit avatar. Sume turns one photo into a reusable avatar handle and finished clips. Pick by who waits on the face.

5 min readSume
All posts

Atmee's LiveKit plugin builds a talking-head avatar from a single portrait (jpg, png or webp) and renders it live in a LiveKit room. Sume also starts from one portrait, but the result is a reusable avatar handle and a finished MP4, not a participant in a call. Choose by whether a person is waiting on the face in real time.

How the Atmee plugin starts

According to the plugin's README, creating an avatar needs a name and one portrait, given as a file path, raw bytes or an HTTPS URL. The avatar is render-ready immediately. You install livekit-plugins-atmee (Python 3.10 or later, livekit-agents 1.6.8 or later), set an ATMEE_API_KEY, and the avatar joins the room as its own participant, lip-synced to your agent's speech.

The README says start() returns in about a second once the render worker acknowledges, and the avatar's tracks appear a few seconds later. The licence is Apache-2.0.

How Sume starts

Sume's Avatar 1.0 has two steps. First, POST /v1/avatar-1.0/generate creates an avatar from a prompt, a profile or a photo (input.type: "photo" with a public HTTPS image_url). The job completes and the avatar becomes a reusable resource in your workspace under a handle. Second, POST /v1/avatar-1.0/talking-video takes that avatar_handle and a script and returns a talking video.

Creating an avatar costs a flat $0.95. A talking video is priced per second by quality: standard $0.184, plus (default) $0.245, max $0.55 per second without a product image. Scripts must estimate to 4 to 60 seconds.

One portrait, two outcomes (read 2026-10-05)
QuestionAtmee pluginSume Avatar 1.0
InputOne portrait, jpg/png/webpOne photo for avatar create, then a script
ResultLive participant in a LiveKit roomReusable handle, then an MP4
Ready timeRender-ready immediately, tracks in a few secondsAvatar create is a job; poll until completed
Cost shapePer minute while the avatar is in the room$0.95 once, then per second of video
60 s of outputMinutes of session time60 x $0.245 = $14.70 at plus quality

Which to pick

If your product is a voice agent that should have a face during a call, the plugin is built for that. If your product is the same presenter appearing in many scheduled videos, a handle is the better fit, since every video reuses the same face. The cost of a brand-consistent presenter walks through twelve videos on one avatar.

Keep the source portrait sharp and front-facing in either case. Both systems can only animate what the photo shows.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume