One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.

Use a single lip-sync model for each assembled video, and measure the frame rate of the first real clip before you pick the output rate. Sume's lip-sync guidance says the model sets the assemble output.fps lock: Fabric is 25 fps, and for MiniMax H3 Max lip sync you must measure it on the first real clip. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session.
Both models take a still image plus audio and return a talking clip. They are different models with different limits, so a video that cuts between them can end up with clips that do not share a frame rate.
The two routes side by side
| Item | VEED Fabric 1.0 | MiniMax H3 Max Lip Sync |
|---|---|---|
| Route | POST /v1/veed/fabric-1.0 | POST /v1/minimax/h3-max/lip-sync |
| Model id | veed/fabric-1.0 | minimax/h3-max/lip-sync |
| Audio window | Up to 300 seconds in the catalog estimate | 5 to 14.8 seconds |
| Resolutions | 480p, 720p | 480p, 768p (default), 1080p; no 2K |
| Rate per audio second | $0.10 at 480p, $0.1875 at 720p | $0.0625 at 480p, $0.10 at 768p, $0.20 at 1080p |
| Default status | Default talk model | Explicit alternative inside the window |
| Frame rate | 25 fps | Measure on the first clip |
Measure it
The check takes seconds. Download the first finished clip from the media URL in the job result and read the stream's frame rate with ffprobe. If it is not what your timeline expects, set your output rate to the one model you chose and render all the clips in the video with that model.
ffprobe -v error -select_streams v:0 \
-show_entries stream=r_frame_rate,width,height \
-of csv=p=0 clip-01.mp4When to pick which
The packet guidance on the lip-sync page says talking shots use text to speech first and then a lip-sync model, never a video model. Fabric stays the default. H3 Max is the explicit alternative when the audio fits the 5 to 14.8 second window. Segments outside the window go to Fabric. If one line in your script is 4 seconds and another is 20, that already decides it: use Fabric for the whole video, so that the rate and look stay uniform.
Do not cut or re-synthesize the audio to make it fit a model's window. The guidance is explicit: never re-cut the audio, because the waveform identity matters. Choose the model that fits the audio, not the other way around.
Cost check
For a 10-second line, H3 Max at 768p is $1.00 and Fabric at 720p is $1.875 (the lip-sync doc reserves $0.94 for five seconds at 720p, which is $0.9375 before rounding). If cost drives the choice and the audio fits, H3 Max is cheaper per second. If a uniform frame rate drives it, choose one model for the whole video even if it costs more.
The workflow in order
- Decide the model from the audio lengths, before you generate any clip.
- Generate the speech with text to speech and measure each file.
- Create the still, inspect it, and fix the face before spending on lip sync.
- Submit the first lip-sync job and wait for it to complete.
- Run ffprobe on the result and note the frame rate.
- Set the assemble output frame rate to match, then submit the remaining lines with the same model.
Through MCP
If an agent drives the work, there is one tool for both talking routes: avatar-image-to-video_create. The model field selects veed/fabric-1.0, which is the default, or minimax/h3-max/lip-sync. The generic generate_video tool refuses the lip-sync id with wrong_tool, so an agent that tries the wrong tool gets a clear next step rather than a wrong clip.
A frame-rate mismatch rarely shows up as an error. The clips render, the job says completed, and the problem appears later as a stutter at a cut, or as audio that slowly drifts from the lips when the timeline resamples one clip but not the other. The lip-sync page sets the rule early to avoid this: one lip-sync model per run, an output rate that matches it, and a measurement on the first real clip so you are not guessing from a spec sheet.
Sources
Related posts
More in Media tools
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
- How to use Sume's Timeline compose and Timeline audio APIs
Timeline compose puts one still and one video in the same frame as a new MP4. Timeline audio joins or splits Sume-hosted audio into reusable files.
Written by Sume