Lip sync audio under 5 seconds: H3 Max rejects it, use Fabric
Sume's MiniMax H3 Max lip sync only accepts 5 to 14.8 seconds of audio and returns invalid_request outside it. A 4.5-second line goes to Fabric instead.

A 4.5-second line sent to Sume's MiniMax H3 Max lip sync is refused with invalid_request: its duration_seconds window is 5 to 14.8, and Sume never clamps a value into it. The fix is to send that clip to VEED Fabric 1.0, which takes the same body and accepts 1 to 300 seconds on its image-to-video route.
The same applies above the window: 15 seconds is refused too.
Why is the window 5 to 14.8 seconds?
It is the provider's. Sume's lip sync note says the upstream rejects audio under 5 seconds and silently clips anything past 14.8 seconds. A clamped reservation would then render a clip shorter than your audio and drop speech, so Sume checks duration_seconds itself and refuses an out-of-range value rather than changing it.
duration_seconds is required and is the basis for the reservation, not a trim instruction. Measure the real audio and send that number.
What does the call look like?
The request is the same for both models: a visual source (image_url or avatar_handle, never both), an audio_url hosted on Sume, duration_seconds, and an optional resolution. Switch the path, not the body: POST /v1/minimax/h3-max/lip-sync for H3 Max and POST /v1/veed/fabric-1.0 for Fabric.
curl -X POST https://api.sume.com/v1/veed/fabric-1.0 \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: short-line-004" \
-d '{
"image_url": "https://example.com/host-still.png",
"audio_url": "https://media.sume.com/example/line-4-5s.mp3",
"duration_seconds": 4.5
}'Which model for which length?
Use the audio length to choose before you submit. Do not pad a short line with silence or cut a long one just to enter the H3 Max window; Sume's guidance is never to re-cut or re-synthesize audio to fit a model.
If you are building an agent, let it measure the duration and route: a clip inside the window may use H3 Max, and everything else goes to Fabric.
| Audio length | Model | What happens |
|---|---|---|
| Under 5 s (for example 4.5 s) | Fabric | Accepted: image-to-video allows 1 to 300 s |
| 5 to 14.8 s | Fabric or H3 Max | Both accepted; H3 Max is the explicit alternative |
| Over 14.8 s (for example 15 s) | Fabric | H3 Max refuses with invalid_request |
What about the 400 and the bill?
An invalid_request is an immediate rejection: Sume only accepts a job when the request can be safely accepted, so nothing is queued. Fix the model or the duration and resubmit. Use a new Idempotency-Key when you change the body, because reusing a key for a different payload returns 409 idempotency_conflict.
If you retry after a timeout, keep the same key with the identical body, and poll the job you already have.
How do I measure the duration reliably?
Probe the file, not the script. A script's spoken length is an estimate; the audio file's real length is a number. Any audio probe gives you the figure to send as duration_seconds; Sume's own Video inspect does the same for a hosted clip.
Send the measured value to the decimal, such as 4.5, not a rounded-down one: duration_seconds is the basis for the price reserved at submit, so an understated number reserves too little and hides the real cost.
Sources
Related posts
More in Sume Avatar 1.0
- Lip sync looks fake? Check teeth, profile, timing and seams on Sume
sync. labs names four tells of fake lip sync. Pull PNG stills from a Sume avatar clip with video frames at the moments each tell shows up, then decide.
- Seasonal avatar host from a prompt, profile or photo for Christmas
Create a reusable Sume Avatar 1.0 handle from a prompt, structured profile props or a reference photo, then reuse it across holiday clips.
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
- MCP avatar-image-to-video_create: choose Fabric or H3 Max
On Sume's hosted MCP, one tool, avatar-image-to-video_create, makes still-plus-audio talking clips. A model field picks Fabric or MiniMax H3 Max lip sync.
Written by Sume