SadTalker is Apache 2.0: self-host it or call a hosted talking clip?

SadTalker turns one portrait and an audio file into a talking head. What it costs you to run, and the hosted Sume still-plus-audio route that skips the GPU.

5 min readSume
All posts

SadTalker is the open-source project built around one idea, "single portrait image + audio = talking head video". Its README shows an Apache 2.0 license, along with a disclaimer against using it for fraud or deception. So the question for a small team is not whether you may run it, but whether you want to run it.

I read the SadTalker README on 2026-10-04. The Sume side comes from the models page and the code behind it.

What running SadTalker involves

The README lists a 256 px and a 512 px face model, a Still mode, and optional GFPGAN enhancement. It asks for Python 3.8 and CUDA 11.3, which you install and keep working yourself.

  • You pin an older Python and CUDA stack and keep it working.
  • You choose the face model size and decide whether to enhance faces.
  • You store and serve the MP4s yourself.

The same inputs on a hosted route

Sume's POST /v1/veed/fabric-1.0 takes the same two things: a still (image_url, or a ready avatar's avatar_handle, never both) and an audio_url hosted on Sume media, up to 10 MB, with a duration_seconds that you measured. It returns a job. A completed job carries a media.sume.com MP4, which you can keep. resolution is 480p or 720p. A 5 second clip at 720p reserves $0.94 and a failed job is refunded.

H3 Max Lip Sync, POST /v1/minimax/h3-max/lip-sync, has the same body but only accepts 5 to 14.8 seconds of audio, with 480p, 768p and 1080p output, and a 5 second clip at 768p reserves $0.50. Pick it when the clip is short and you want 1080p.

A quick way to choose

Choose self-hosting when you already own the GPU, want the face model under your control, and can absorb the maintenance. Choose the hosted route when you have a few dozen clips a month and no GPU, because you pay per second of audio and nothing sits idle. If the audio does not exist yet, one request to POST /v1/avatar-1.0/talking-video takes a script instead, and runs 4 to 60 seconds.

Either way, check the photo's rights first. The SadTalker disclaimer and the same caution applies to any avatar tool: animate people who have agreed to it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume