Alibaba Live Avatar needs five H800 GPUs: or call a hosted clip API
Alibaba-Quark's open Live Avatar streams audio-driven video at 45 FPS on five H800 GPUs. Compare that hardware with a hosted still-plus-audio clip.
Alibaba-Quark's Live Avatar is an Apache 2.0 licensed, 14-billion-parameter model that reaches 45 FPS on multiple H800 GPUs, but real-time inference needs five of them with 80 GB each. If you only need finished talking clips, a hosted lip-sync call turns a still and audio into an MP4 at $0.10 a second at 768p on Sume, with no GPU to run.
What the repository says
The GitHub page describes Live Avatar as streaming, audio-driven avatar generation of infinite length. It is built as a LoRA adapter on the WanS2V-14B base model, uses 4-step sampling, and handles videos beyond 10,000 seconds with block-wise autoregressive processing.
For hardware, the page lists five H800 GPUs (80 GB VRAM each) for real-time inference with the TPP pipeline. A single GPU needs at least 80 GB of VRAM, and an FP8 mode can run on 48 GB cards with a slight quality trade-off. Weights landed on December 8, 2025, and the project was accepted as an ECCV 2026 Spotlight on June 18, 2026.
Self-host or call an API
Running it yourself makes sense when you need an always-on stream, own the GPUs, and want the licence freedom. It is the wrong tool when the job is a 10-second greeting that renders once.
On Sume you send image_url or an avatar handle, a Sume-hosted audio_url of 5 to 14.8 seconds, and a duration_seconds value. The API reserves ceil(duration_seconds) times the per-second rate and refunds it if the job fails.
| Item | Live Avatar (self-hosted) | Sume H3 Max lip sync |
|---|---|---|
| Licence | Apache 2.0 | Hosted API, no weights |
| Hardware | 5 x 80 GB H800 for real-time; 80 GB minimum on one GPU | None on your side |
| Length | Infinite, over 10,000 s shown | 5 to 14.8 s per clip |
| Output | Live stream | MP4 job result |
| Price shape | Your GPU hours | 768p: $0.08 x 1.25 = $0.10 per second |
| Example | One hour of H800 time, your rate | 15 s at 768p: 15 x $0.10 = $1.50 |
A fair test before you commit
Take one real 12-second script. Render it with the open model on a rented GPU and with the hosted route, and compare three things: the total cost including setup time, whether the mouth shapes match the audio, and how much glue code you wrote.
Check the mouth by reading the transcript word times against frames, as in checking avatar lip sync by transcript word times. The open model's strength is length. The hosted clip's strength is that the bill is a single line per clip.
Sources
Related posts
More in Comparisons
- Anam plans from $12 to $999: cost per live minute vs a rendered one
Anam's plan ladder works out to $0.12 to $0.24 per included live minute. A 60-second Sume avatar clip costs $11.04, so it wins above about 74 viewers.
- Atmee avatar billed by the minute vs Sume reserve and refund
Atmee's LiveKit plugin bills by the minute, capped at 3,600 seconds by default. Sume reserves a clip price upfront and refunds failures. Compared.
- Atmee LiveKit avatar from one portrait vs a Sume photo avatar clip
Atmee turns one portrait into a live LiveKit avatar. Sume turns one photo into a reusable avatar handle and finished clips. Pick by who waits on the face.
- Best API for text to speech: six checks to run before you pick
Compare TTS APIs on price per million characters, request cap, latency, languages, controls and word timing, with MAI-Voice-2.1 and Sume TTS 1.0 numbers.
Written by Sume