Luma API concurrent jobs limit vs Sume plan concurrency
Luma caps active generations per API client and answers 429 when full. Sume ties concurrency to your plan and queues extra jobs until queue_full.

Luma's concurrent-jobs limit is a ceiling on active, non-terminal generations per API client, and exceeding it returns HTTP 429 just like the per-minute limit. Sume's concurrency is set by your plan alone, and a full processing slot is not an error: valid jobs wait as queued until the queue is also full, which returns 429 queue_full.
What does Luma's concurrent-jobs limit cover?
Luma's rate-limit guide, read 2026-10-01, names two independent limits on POST /v1/generations: requests per minute and concurrent jobs, the maximum number of active (non-terminal) generations at any time. Both are evaluated per API client and a request must pass both. The numbers depend on your plan and appear in the Luma platform dashboard.
How does Sume set concurrency?
The generation admission page says generation concurrency is plan-only: prepaid top-ups do not raise it, and only admin overrides can. Queue capacity defaults to max(3, concurrency_limit × 5). The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit, so prefer that field over a static table.
| Plan | Processing | Queue (default) | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What happens at the limit on each side?
On Luma, a full concurrent-jobs ceiling refuses the new request. On Sume, with concurrency_limit: 1 you may submit several valid jobs at once and all can come back queued; only one same-workspace job moves to processing at a time. The submit fails with 429 queue_full only when "the workspace has no remaining accepted generation capacity".
What should my client do?
Do not treat queued as failure. Store the job_id, poll with backoff, and fetch the result when status is completed. On queue_full, wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. See queue_full versus a full concurrency slot for the distinction.
Sources
Related posts
More in Developers
- Luma API video URL expires after 1 hour: what to do on Sume
Luma Agents API video URLs are presigned and expire after 1 hour. On Sume, a completed job is fetched with your API key from the /content endpoint.
- Luma API X-Request-Id vs Sume x-sume-request-id
Luma's X-Request-Id echoes your header or is generated. Sume sends x-sume-request-id on every response; quote it, plus error.code, when you contact support.
- Luma Ray 3.2 extend with a generation id vs Sume per-clip jobs
Luma extends video from one prior generation id as start or end frame. Sume docs describe frame_images with first_frame or last_frame, a new job per clip.
- Ray 3.2 API for render farms: the Sume callback_url pattern
Luma pitches the Ray3.2 API for pipelines and render farms. On Sume, wire an async video job with callback_url, Idempotency-Key and a request id to log.
Written by Sume