Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?
Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.

If you already have a voice recording and a face still, Sume offers two routes: VEED Fabric 1.0 (veed/fabric-1.0), which the docs list for 1 to 300 seconds, and MiniMax H3 Max Lip Sync, which takes the same still-plus-audio body with audio of 5 to 14.8 seconds. Use Fabric for anything longer than about 15 seconds, and Avatar video when you want Sume to voice a script.
Side by side
From the Models overview, read 2026-10-08:
| Route | Input | Duration |
|---|---|---|
| Avatar 1.0 talking-video | avatar_handle plus script or video_inputs | Estimated 4-60 s |
| VEED Fabric 1.0 | audio_url, duration_seconds and one of image_url or avatar_handle | 1-300 s on the image-to-video body |
| MiniMax H3 Max Lip Sync | Same still plus audio body as Fabric | Audio 5-14.8 s |
Fabric request rules
Send audio_url, a measured duration_seconds and exactly one visual source. The preferred source is the image_url of a generated, inspected posed still. Use avatar_handle only when the user named that avatar. The two sources cannot be sent together. Measure the duration yourself rather than guessing, because it sets what the job covers.
- Visual source: image_url or avatar_handle, never both.
- Audio: a public HTTPS audio_url.
- Aliases: /v1/avatar-1.0/image-to-video is deprecated; use /v1/veed/fabric-1.0.
- The experimental /v1/avatar-1.0/fabric route is test-only and the docs say not to build production integrations on it.
How to choose
Start from what you hold. If you only have text, use Avatar video and let Sume voice the script. If you have a recorded read of 15 seconds or less and want the best match to your voice, either lip-sync route can work, and you can compare outputs on the same inputs. If your recording runs longer, use Fabric as one job rather than cutting it into 14.8-second pieces; the linked post shows that arithmetic.
Remember the rule behind all of this: a video model clip with narration laid over it will not lip-sync, so speech shots belong on one of these routes.
Mistakes to avoid
Do not send both image_url and avatar_handle to Fabric; the request is rejected. Do not guess duration_seconds; measure the audio. Do not feed audio longer than 14.8 seconds to the H3 Max route. Do not build on the experimental fabric route, since its name will change.
Check the still first. A posed, inspected frame with a clear face gives the model a better start than a crop with the mouth hidden. Keep the audio clean, with little room echo, because the lip match can only be as good as the signal it hears.
Sources
Related posts
More in Sume Avatar 1.0
- Reuse one AI spokesperson across videos with an avatar handle
A Sume avatar_handle is a stable name for a ready avatar. Sume strips a leading @, so you create once and call the handle in every talking-video request.
- Add a silent beat to an AI avatar video with voice type silence
In Sume multi-scene avatar videos a scene with voice type silence is a pause with no speech. It needs a duration and rejects script text. Rules and example.
- Sume Avatar 1.0 ratios vs TikTok ad ratios: three overlap, two do not
Avatar 1.0 makes 1:1, 3:4, 9:16, 4:3 or 16:9 at 720p, 4 to 60 s. TikTok's non-Spark ad page lists 9:16, 16:9 and 1:1, so 3:4 and 4:3 need a crop.
- Does Sume Avatar 1.0 speak Spanish or Korean? English only in code
The Avatar 1.0 talking-video prompt on main says English only. What that means for Spanish or Korean lines, and the audio-driven route to test instead.
Written by Sume