Alibaba Live Avatar needs five H800 GPUs: or call a hosted clip API

Alibaba-Quark's open Live Avatar streams audio-driven video at 45 FPS on five H800 GPUs. Compare that hardware with a hosted still-plus-audio clip.

5 min readSume
All posts

Alibaba-Quark's Live Avatar is an Apache 2.0 licensed, 14-billion-parameter model that reaches 45 FPS on multiple H800 GPUs, but real-time inference needs five of them with 80 GB each. If you only need finished talking clips, a hosted lip-sync call turns a still and audio into an MP4 at $0.10 a second at 768p on Sume, with no GPU to run.

What the repository says

The GitHub page describes Live Avatar as streaming, audio-driven avatar generation of infinite length. It is built as a LoRA adapter on the WanS2V-14B base model, uses 4-step sampling, and handles videos beyond 10,000 seconds with block-wise autoregressive processing.

For hardware, the page lists five H800 GPUs (80 GB VRAM each) for real-time inference with the TPP pipeline. A single GPU needs at least 80 GB of VRAM, and an FP8 mode can run on 48 GB cards with a slight quality trade-off. Weights landed on December 8, 2025, and the project was accepted as an ECCV 2026 Spotlight on June 18, 2026.

Self-host or call an API

Running it yourself makes sense when you need an always-on stream, own the GPUs, and want the licence freedom. It is the wrong tool when the job is a 10-second greeting that renders once.

On Sume you send image_url or an avatar handle, a Sume-hosted audio_url of 5 to 14.8 seconds, and a duration_seconds value. The API reserves ceil(duration_seconds) times the per-second rate and refunds it if the job fails.

Self-hosted Live Avatar vs a hosted clip (read 2026-10-05)
ItemLive Avatar (self-hosted)Sume H3 Max lip sync
LicenceApache 2.0Hosted API, no weights
Hardware5 x 80 GB H800 for real-time; 80 GB minimum on one GPUNone on your side
LengthInfinite, over 10,000 s shown5 to 14.8 s per clip
OutputLive streamMP4 job result
Price shapeYour GPU hours768p: $0.08 x 1.25 = $0.10 per second
ExampleOne hour of H800 time, your rate15 s at 768p: 15 x $0.10 = $1.50

A fair test before you commit

Take one real 12-second script. Render it with the open model on a rented GPU and with the hosted route, and compare three things: the total cost including setup time, whether the mouth shapes match the audio, and how much glue code you wrote.

Check the mouth by reading the transcript word times against frames, as in checking avatar lip sync by transcript word times. The open model's strength is length. The hosted clip's strength is that the bill is a single line per clip.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume