MiniMax H3 Max video or H3 Max lip sync: fixing the wrong_tool error
generate_video refuses minimax/h3-max/lip-sync with wrong_tool. Use avatar-image-to-video_create for lip sync, minimax-h3-max for text or image video.

If generate_video returns wrong_tool for minimax/h3-max/lip-sync, you called the video tool for a lip-sync job. Lip sync is job type avatar_image_to_video, and its MCP tool is avatar-image-to-video_create. The video model minimax-h3-max is a different catalog entry that makes video from a prompt or image, with no audio you supply. MiniMax H3 was recorded as released on 2026-07-31 with open weights on 2026-08-03 (Magic Hour tracker, read 2026-10-06); Sume's ids are the ones in Models.
Which id goes to which tool?
| You want | Model id | Tool or route | Job type |
|---|---|---|---|
| Video from a prompt or image | minimax-h3-max | generate_video / POST /v1/videos | video generation job |
| A still that says your audio | minimax/h3-max/lip-sync | avatar-image-to-video_create / POST /v1/minimax/h3-max/lip-sync | avatar_image_to_video |
| A still that copies a motion video | kling/3.0/motion-control | POST /v1/kling/3.0/motion-control | avatar_image_to_video |
| A talking default | veed/fabric-1.0 | Fabric talk route | avatar_image_to_video |
What does the correct call look like?
Replace the two media URLs with your own Sume-hosted files. Do not put model, endpoint or provider_endpoint in the body; the route rejects them because the path already names the model.
import os, requests
BASE = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "lipsync-demo-001"}
body = {
"image_url": "https://media.sume.com/demo/presenter.png",
"audio_url": "https://media.sume.com/demo/line.mp3",
"duration_seconds": 8,
"resolution": "768p",
}
r = requests.post(BASE + "/v1/minimax/h3-max/lip-sync",
json=body, headers=HEAD, timeout=30)
print(r.status_code, r.json())
Why the two ids are kept apart
The split is deliberate. The lip-sync route is described in Sume's API notes as the fourth model on the avatar_image_to_video job type, alongside Fabric and Kling motion control, and the public model id rather than a new job type separates them. The video model minimax-h3-max sits in the Video Router catalog instead, with prompt-driven text, image and reference modes at 480p, 768p and 1080p for 5 to 15 seconds and no 2K or 4K. Because each surface has its own strict body, a request shaped for one is not silently reinterpreted by the other.
The same strictness applies to the lip-sync body: it rejects model, endpoint and provider_endpoint, so the URL is the model. If you copy a generate_video payload that carries model: "minimax/h3-max/lip-sync", that is exactly the mistake the wrong_tool error is there to catch, and the MCP error names the next action, call_avatar_image_to_video_create.
On cost, the two ids bill differently even though both carry the 1.25 house margin. The video model's list rates are $0.05, $0.08 and $0.16 a second at 480p, 768p and 1080p, and the lip-sync route uses the same three list numbers on its own rate card, so a 5 second 768p clip reserves $0.50 on the lip-sync route. Check the Billing section of each route before you compare totals.
What else trips people?
The wrong-tool post gives the same fix for motion control, and the MiniMax price table lists what each id costs. The error is cheap: it is raised before any provider work, so nothing is reserved when you hit it.
- duration_seconds is required and must be 5 to 14.8. It is the reservation basis and is never clamped.
- Send exactly one of image_url, avatar_id or avatar_handle.
- At 768p a 5 second clip reserves $0.50 and is captured on completion; a failed job is refunded.
- A 2K request is invalid. Resolutions are 480p, 768p and 1080p.
How do you stop the error returning?
Route by model id in your code, not by habit. Keep a small map from id to tool or path, so minimax/h3-max/lip-sync always goes to the avatar route and minimax-h3-max to the video route. Add a test that fails if an id is missing from the map, and log the wrong_tool code with the id so a misrouted call is easy to find.
If you work through MCP, let the agent read the error: it names the right tool as the next action. The fix is a single changed tool name, with the same arguments mapped to the lip-sync body, plus the required duration_seconds.
Sources
Related posts
More in Media tools
- Reddit's 15-second view rule: a lip-sync clip that stays under it
Reddit's Engaged Video Views beta bills videos over 15 s at 15 s. Sume H3 Max lip sync tops out at 14.8 s of audio, so one clip fits. Specs, math and a caveat.
- Reel captions in two languages: language hint vs auto-detect
How Sume video captions pick a language and a style: the optional language hint, auto-detect, the Hangul rule that returns 400, and a Korean Reel example.
- Reels profile grid crops to 3:4 (1080x1440): check a frame
A Reel cover is center-cropped to 3:4 on the profile grid, per a third-party guide. Pull the frame with Sume video frames and keep faces and titles inside.
- Shorts episode title card: a 2-second still, then the clip
Open each Shorts episode on a held title still with Timeline 1.0: slot math, the 0.5 s coverage rule, what stills ignore, and one render body you can copy.
Written by Sume