LivePortrait needs a driving video, not audio: Kling motion control
LivePortrait animates a portrait from a driving video, not from speech. Sume's Kling 3.0 motion control route does that as a hosted API, per second.

If you searched for LivePortrait expecting a tool that makes a photo talk from an audio file, note what it takes: a source portrait plus a driving video, whose head and face motion are copied onto the portrait. The README for the KwaiVGI project describes it that way, with a humans mode and an animals mode for cats and dogs. It uses InsightFace for face detection and carries its own misuse disclaimer. I read it on 2026-10-04.
Sume has a route with the same input shape: Kling 3.0 motion control.
Same idea, hosted
POST /v1/kling/3.0/motion-control, also reachable at the alias /v1/avatar-1.0/motion-control, takes a still image and a motion_video_url. The still is the character, and the driving clip supplies the movement. Duration runs from 1 to 30 seconds. The list price is $0.126 per second, and Sume bills 1.25 times list, so a 10 second clip costs 10 x $0.126 x 1.25 = $1.575.
The job is asynchronous like the other video routes: you receive a job id, poll it or take a webhook, and the finished MP4 is on a Sume media URL.
| Question | LivePortrait | Sume Kling motion control |
|---|---|---|
| Drives the face with | A driving video | A driving video (motion_video_url) |
| Audio-driven speech | No, motion only | No, motion only |
| Subjects | Humans and animals modes | A still image you supply |
| Length | Set by your driving clip | 1 to 30 s |
| Runs on | Your machine | Sume job, per-second billing |
If you want speech, use a different route
Neither tool turns text into a talking face by itself. For that, Sume has POST /v1/avatar-1.0/talking-video (a script and a ready avatar, 4 to 60 seconds) and the audio-driven lip-sync routes Fabric and H3 Max. Use motion control when you want to borrow a performance: a nod, a head turn, a gesture you recorded on your phone. Use the talking-video route when you start from words.
A common pairing is to record a short driving clip of yourself, apply it to a brand character with motion control, and add the voice separately in editing. Only animate faces you have the rights to use.
Sources
Related posts
More in Sume Avatar 1.0
- Nonprofit donor thank-you: three tiered avatar videos, total cost
Three 20-second thank-you videos, one per giving tier, cost $11.04 in total at standard. See how to script them and why you review a preview first.
- Audio2Face-3D blendshapes vs a finished lip-synced MP4 from a photo
NVIDIA Audio2Face-3D turns audio into facial animation data for a 3D rig. If you want a finished MP4 from a photo instead, this is the Sume route and its cost.
- Podcast episode promo: a 20-second avatar host clip, cost by tier
Turn an episode summary into a 20-second vertical promo with an AI avatar host. Standard costs $3.68, plus $4.90, max $11.00. Includes captions and a script.
- SadTalker is Apache 2.0: self-host it or call a hosted talking clip?
SadTalker turns one portrait and an audio file into a talking head. What it costs you to run, and the hosted Sume still-plus-audio route that skips the GPU.
Written by Sume