Audio2Face-3D blendshapes vs a finished lip-synced MP4 from a photo
NVIDIA Audio2Face-3D turns audio into facial animation data for a 3D rig. If you want a finished MP4 from a photo instead, this is the Sume route and its cost.

Audio2Face-3D is NVIDIA's technology for generating facial animation from audio. Per its README, it takes audio in and produces mesh deformation, joint transforms or blendshape weights as output, and it runs in real time or in batch. Those are inputs to a 3D character pipeline, not a video file. The README mentions Maya and Unreal Engine 5 plugins for that purpose.
So the buyer question splits in two: do you have a 3D character that needs animating, or do you want a finished clip of a face speaking? I read the README on 2026-10-04 to separate the two.
What comes out
| Question | Audio2Face-3D | Sume avatar video |
|---|---|---|
| Input | Audio | A script, or voiced scenes, plus an avatar |
| Output | Blendshape weights, joint transforms or mesh deformation | A 720p MP4 |
| Needs a 3D rig | Yes | No, one avatar image |
| Licensing | SDK MIT; training Apache; models under NVIDIA licenses | Per-second API pricing |
| Rendering | You render it in your engine | Rendered by Sume |
A note on licenses
The README splits terms by part: the SDK is MIT, the training framework is Apache, and the models sit under an NVIDIA open model license or an evaluation-only license. If you plan to ship a product with it, read which license covers the model you use.
The photo-to-MP4 route
If your goal is a presenter video and you have no rig, POST /v1/avatar-1.0/talking-video gives you one request with a script. You choose the quality (standard $0.184, plus $0.245 or max $0.55 per second without a product image), an aspect ratio, and optional burned-in captions. A 20 second clip at standard costs 20 x $0.184 = $3.68.
You make the avatar once with POST /v1/avatar-1.0/generate ($0.95 flat), from a prompt, a profile, or a public photo. Then reuse its handle in every video.
Which one for a game or a film pipeline?
For a game, a virtual human in an engine, or a film shot, blendshape output is what you want, and Sume does not produce it. For ads, explainers and social posts, the rendered MP4 saves the entire rig step. The two are not rivals; they sit on different ends of the same task.
Sources
Related posts
More in Sume Avatar 1.0
- Podcast episode promo: a 20-second avatar host clip, cost by tier
Turn an episode summary into a 20-second vertical promo with an AI avatar host. Standard costs $3.68, plus $4.90, max $11.00. Includes captions and a script.
- SadTalker is Apache 2.0: self-host it or call a hosted talking clip?
SadTalker turns one portrait and an audio file into a talking head. What it costs you to run, and the hosted Sume still-plus-audio route that skips the GPU.
- School principal weekly update: a 40-second avatar clip, 36 weeks
A 40-second weekly announcement from one AI avatar costs $7.36 at standard, or $264.96 for a 36-week school year. Includes a script frame and caption tip.
- Setup guide intro with a product image: +$0.01 to $0.03 a second
Adding product_image to a Sume avatar video raises the per-second rate by $0.01 on Standard, $0.013 on Plus and $0.03 on Max: +$0.30, $0.39 or $0.90 for 30 s.
Written by Sume