MiniMax H3 API rate limit: 300 RPM, 30 in-flight vs Sume
MiniMax lists 300 requests per minute and 30 in-flight tasks for H3, and 20 RPM for Hailuo. Sume queues by plan concurrency instead. Read both limits here.

MiniMax's rate-limit page lists 300 requests per minute and 30 in-flight tasks for the H3 video endpoint (MiniMax-H3), and 20 requests per minute for the older Hailuo series, with no in-flight cap given (read 2026-10-02). On Sume you do not tune a per-model number: jobs are accepted into a queue and a plan-based concurrency cap decides how many run at once.
The two limits answer different questions. MiniMax's RPM caps how fast you may submit; the in-flight cap bounds how many H3 tasks can be running or waiting at your account. Sume's concurrency caps how many of your workspace's paid jobs are processing together.
What does MiniMax list for video?
The page lists two video rows. It does not list a separate row for MiniMax-H3-Max, so check the dashboard before assuming H3 Max shares H3's numbers.
| Video endpoint | Requests per minute | Max in-flight tasks |
|---|---|---|
| MiniMax-H3 (video generation V2) | 300 | 30 |
| Hailuo series (video generation) | 20 | Not listed |
What does Sume do instead?
Sume's generation admission docs describe queue-first handling. A valid job can be accepted as queued even when your workspace is at its concurrency limit, and workers move it to processing when a slot opens. Concurrency is plan-only; prepaid top-ups do not raise it.
Three separate errors can still reach you: 429 queue_full when the queue is out of room, 429 rate_limited for submit-request volume, and 402 insufficient_credits when the balance cannot cover the reservation. Retry the first two with backoff and the same Idempotency-Key, so a replay returns the original job instead of a second charge.
| Plan | Processing concurrency | Queue capacity (default) |
|---|---|---|
| Free | 1 | 5 |
| Pro | 4 | 20 |
| Startup | 8 | 40 |
| Scale | 20 | 100 |
| Enterprise | 20 | 100 |
How should you size a batch?
Against MiniMax directly, a batch of 100 H3 clips needs a client that never holds more than 30 tasks open and stays under 300 submissions a minute. MiniMax's guide recommends polling every 10 seconds, so 30 open tasks is about 3 status calls a second at worst (read 2026-10-02).
On Sume, submit the whole batch with minimax-h3 or minimax-h3-max and let the queue drain. Read generation_limits.concurrency_limit from your workspace rather than the static table above, because admin overrides can raise it. Sume does not publish a per-model RPM for these ids, so do not hard-code one.
What this does not tell you
The MiniMax page gives no burst or retry-after rules, and the Hailuo row says nothing about in-flight tasks. Sume's docs likewise state no per-model submit rate for minimax-h3. For production sizing, test with a small paid batch and read the 429 responses you actually get.
Sources
Related posts
More in Models
- MiniMax H3 camera prompts: lens, movement, exposure wording
fal's H3 prompting guide says H3 reads film vocabulary: lens, rack focus, handheld, grain. Wording that works as a prompt and a Sume request that sends it.
- MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.
- MiniMax H3 reference inputs: 12 files combined, not 9 + 3 + 3
MiniMax H3 allows 9 images, 3 videos and 3 audio clips as references, but a 12-file combined cap applies. How to stay inside it, and how Sume counts.
- MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.
Written by Sume