Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?

Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.

5 min readSume
All posts

If you already have a voice recording and a face still, Sume offers two routes: VEED Fabric 1.0 (veed/fabric-1.0), which the docs list for 1 to 300 seconds, and MiniMax H3 Max Lip Sync, which takes the same still-plus-audio body with audio of 5 to 14.8 seconds. Use Fabric for anything longer than about 15 seconds, and Avatar video when you want Sume to voice a script.

Side by side

From the Models overview, read 2026-10-08:

Talking-face routes, as of 2026-10-08
RouteInputDuration
Avatar 1.0 talking-videoavatar_handle plus script or video_inputsEstimated 4-60 s
VEED Fabric 1.0audio_url, duration_seconds and one of image_url or avatar_handle1-300 s on the image-to-video body
MiniMax H3 Max Lip SyncSame still plus audio body as FabricAudio 5-14.8 s

Fabric request rules

Send audio_url, a measured duration_seconds and exactly one visual source. The preferred source is the image_url of a generated, inspected posed still. Use avatar_handle only when the user named that avatar. The two sources cannot be sent together. Measure the duration yourself rather than guessing, because it sets what the job covers.

  • Visual source: image_url or avatar_handle, never both.
  • Audio: a public HTTPS audio_url.
  • Aliases: /v1/avatar-1.0/image-to-video is deprecated; use /v1/veed/fabric-1.0.
  • The experimental /v1/avatar-1.0/fabric route is test-only and the docs say not to build production integrations on it.

How to choose

Start from what you hold. If you only have text, use Avatar video and let Sume voice the script. If you have a recorded read of 15 seconds or less and want the best match to your voice, either lip-sync route can work, and you can compare outputs on the same inputs. If your recording runs longer, use Fabric as one job rather than cutting it into 14.8-second pieces; the linked post shows that arithmetic.

Remember the rule behind all of this: a video model clip with narration laid over it will not lip-sync, so speech shots belong on one of these routes.

Mistakes to avoid

Do not send both image_url and avatar_handle to Fabric; the request is rejected. Do not guess duration_seconds; measure the audio. Do not feed audio longer than 14.8 seconds to the H3 Max route. Do not build on the experimental fabric route, since its name will change.

Check the still first. A posed, inspected frame with a clear face gives the model a better start than a crop with the mouth hidden. Keep the audio clean, with little room echo, because the lip match can only be as good as the signal it hears.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume