Audio2Face-3D blendshapes vs a finished lip-synced MP4 from a photo

NVIDIA Audio2Face-3D turns audio into facial animation data for a 3D rig. If you want a finished MP4 from a photo instead, this is the Sume route and its cost.

5 min readSume
All posts

Audio2Face-3D is NVIDIA's technology for generating facial animation from audio. Per its README, it takes audio in and produces mesh deformation, joint transforms or blendshape weights as output, and it runs in real time or in batch. Those are inputs to a 3D character pipeline, not a video file. The README mentions Maya and Unreal Engine 5 plugins for that purpose.

So the buyer question splits in two: do you have a 3D character that needs animating, or do you want a finished clip of a face speaking? I read the README on 2026-10-04 to separate the two.

What comes out

Audio2Face-3D README vs Sume avatar video (read 2026-10-04)
QuestionAudio2Face-3DSume avatar video
InputAudioA script, or voiced scenes, plus an avatar
OutputBlendshape weights, joint transforms or mesh deformationA 720p MP4
Needs a 3D rigYesNo, one avatar image
LicensingSDK MIT; training Apache; models under NVIDIA licensesPer-second API pricing
RenderingYou render it in your engineRendered by Sume

A note on licenses

The README splits terms by part: the SDK is MIT, the training framework is Apache, and the models sit under an NVIDIA open model license or an evaluation-only license. If you plan to ship a product with it, read which license covers the model you use.

The photo-to-MP4 route

If your goal is a presenter video and you have no rig, POST /v1/avatar-1.0/talking-video gives you one request with a script. You choose the quality (standard $0.184, plus $0.245 or max $0.55 per second without a product image), an aspect ratio, and optional burned-in captions. A 20 second clip at standard costs 20 x $0.184 = $3.68.

You make the avatar once with POST /v1/avatar-1.0/generate ($0.95 flat), from a prompt, a profile, or a public photo. Then reuse its handle in every video.

Which one for a game or a film pipeline?

For a game, a virtual human in an engine, or a film shot, blendshape output is what you want, and Sume does not produce it. For ads, explainers and social posts, the rendered MP4 saves the entire rig step. The two are not rivals; they sit on different ends of the same task.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume