Luma API concurrent jobs limit vs Sume plan concurrency

Luma caps active generations per API client and answers 429 when full. Sume ties concurrency to your plan and queues extra jobs until queue_full.

4 min readSume
All posts

Luma's concurrent-jobs limit is a ceiling on active, non-terminal generations per API client, and exceeding it returns HTTP 429 just like the per-minute limit. Sume's concurrency is set by your plan alone, and a full processing slot is not an error: valid jobs wait as queued until the queue is also full, which returns 429 queue_full.

What does Luma's concurrent-jobs limit cover?

Luma's rate-limit guide, read 2026-10-01, names two independent limits on POST /v1/generations: requests per minute and concurrent jobs, the maximum number of active (non-terminal) generations at any time. Both are evaluated per API client and a request must pass both. The numbers depend on your plan and appear in the Luma platform dashboard.

How does Sume set concurrency?

The generation admission page says generation concurrency is plan-only: prepaid top-ups do not raise it, and only admin overrides can. Queue capacity defaults to max(3, concurrency_limit × 5). The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit, so prefer that field over a static table.

Sume plan limits from the admission docs, read 2026-10-01.
PlanProcessingQueue (default)Accepted jobs
Free156
Pro42024
Startup84048
Scale20100120

What happens at the limit on each side?

On Luma, a full concurrent-jobs ceiling refuses the new request. On Sume, with concurrency_limit: 1 you may submit several valid jobs at once and all can come back queued; only one same-workspace job moves to processing at a time. The submit fails with 429 queue_full only when "the workspace has no remaining accepted generation capacity".

What should my client do?

Do not treat queued as failure. Store the job_id, poll with backoff, and fetch the result when status is completed. On queue_full, wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. See queue_full versus a full concurrency slot for the distinction.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume