MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.

fal lists MiniMax H3 Max as a post-trained variant of MiniMax H3 tuned for stronger prompt adherence. On Sume it is minimax-h3-max: text-to-video, start and end frame image-to-video, and reference-to-video at 480p, 768p or 1080p for 5 to 15 seconds, with native stereo audio. Whether it follows your prompt better than plain H3 is something to test on your own prompts.
The two descriptions
The first column is from fal's model explorer. The Sume column is from the Video Router and video generation docs. The two sources describe the model differently, one by tuning and one by speed, so neither tells you the whole story.
| Item | fal listing | Sume docs |
|---|---|---|
| What it is | Post-trained variant of MiniMax H3, tuned for stronger prompt adherence | The faster 768p variant |
| Modes | Also lists H3 Max image-to-video and H3 (2K image-to-video) | Text-to-video, first and last frame image-to-video, reference-to-video |
| Resolution | Not covered here | 480p, 768p, 1080p (1080p is a latent refinement from native 768p) |
| Duration | Not covered here | 5 to 15 s |
| Audio | Not covered here | Native stereo audio |
Call it on Sume
Billing is provider list times 1.25 like every Sume video model, and the catalog carries the numbers. The model reports image, video and audio supported_input_references. The call below is text-to-video at 768p on the Video Router route, which Sume says stays available; new integrations are pointed to POST /v1/videos, where the same model id works.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-max-001" \
-d '{
"model": "minimax-h3-max",
"prompt": "A slow dolly-in on a copper kettle as steam rises, morning light",
"resolution": "768p",
"duration": 8,
"aspect_ratio": "16:9",
"mode": "async"
}'A fair prompt-adherence test
- Write five prompts with three concrete, checkable details each (object, motion, light).
- Run each on minimax-h3 and minimax-h3-max at the same resolution and duration.
- Score each clip by counting how many of the three details appear.
- Read capabilities from GET /v1/video-router/models rather than assuming one envelope.
Sources
Related posts
More in Models
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- Nano Banana reference image limits: Lite, 2 and Pro compared
Nano Banana 2 Lite takes up to 14 object images, Nano Banana 2 takes 10 object, 4 character and 3 style, Pro takes 6 object and 5 character.
- Pixal3D multi-view: prepare the input images
Pixal3D added multi-view inference in September 2026 under an MIT license. Sume has no 3D, but reference-image edit can prepare input views.
- PixVerse V6 native audio and camera work: what to check in an API
PixVerse's blog lists V6 with camera work and native audio, plus R2 and a $439M Series C total. How to test those claims against any video API's catalog.
Written by Sume