MiniMax H3 API rate limit: 300 RPM, 30 in-flight vs Sume

MiniMax lists 300 requests per minute and 30 in-flight tasks for H3, and 20 RPM for Hailuo. Sume queues by plan concurrency instead. Read both limits here.

5 min readSume
All posts

MiniMax's rate-limit page lists 300 requests per minute and 30 in-flight tasks for the H3 video endpoint (MiniMax-H3), and 20 requests per minute for the older Hailuo series, with no in-flight cap given (read 2026-10-02). On Sume you do not tune a per-model number: jobs are accepted into a queue and a plan-based concurrency cap decides how many run at once.

The two limits answer different questions. MiniMax's RPM caps how fast you may submit; the in-flight cap bounds how many H3 tasks can be running or waiting at your account. Sume's concurrency caps how many of your workspace's paid jobs are processing together.

What does MiniMax list for video?

The page lists two video rows. It does not list a separate row for MiniMax-H3-Max, so check the dashboard before assuming H3 Max shares H3's numbers.

MiniMax video rate limits (read 2026-10-02)
Video endpointRequests per minuteMax in-flight tasks
MiniMax-H3 (video generation V2)30030
Hailuo series (video generation)20Not listed

What does Sume do instead?

Sume's generation admission docs describe queue-first handling. A valid job can be accepted as queued even when your workspace is at its concurrency limit, and workers move it to processing when a slot opens. Concurrency is plan-only; prepaid top-ups do not raise it.

Three separate errors can still reach you: 429 queue_full when the queue is out of room, 429 rate_limited for submit-request volume, and 402 insufficient_credits when the balance cannot cover the reservation. Retry the first two with backoff and the same Idempotency-Key, so a replay returns the original job instead of a second charge.

Sume processing concurrency by plan (Sume docs)
PlanProcessing concurrencyQueue capacity (default)
Free15
Pro420
Startup840
Scale20100
Enterprise20100

How should you size a batch?

Against MiniMax directly, a batch of 100 H3 clips needs a client that never holds more than 30 tasks open and stays under 300 submissions a minute. MiniMax's guide recommends polling every 10 seconds, so 30 open tasks is about 3 status calls a second at worst (read 2026-10-02).

On Sume, submit the whole batch with minimax-h3 or minimax-h3-max and let the queue drain. Read generation_limits.concurrency_limit from your workspace rather than the static table above, because admin overrides can raise it. Sume does not publish a per-model RPM for these ids, so do not hard-code one.

What this does not tell you

The MiniMax page gives no burst or retry-after rules, and the Hailuo row says nothing about in-flight tasks. Sume's docs likewise state no per-model submit rate for minimax-h3. For production sizing, test with a small paid batch and read the 429 responses you actually get.

Sources

Related posts

More in Models

All Models posts

Written by Sume