Luma Build tier: 10 concurrent Ray jobs, 20 requests a minute
Luma's Build tier allows 10 concurrent Ray video jobs, 20 requests a minute and $5000 a month. How Sume's plan concurrency and queue capacity differ.

On Luma's documented Build tier, the Ray video model allows 10 concurrent generations and 20 requests per minute, the Photon image models allow 40 and 80, and the tier has a usage limit of $5000 per month. Sume describes limits differently: generation concurrency is set by your plan, queue capacity sits on top of it, and prepaid top-ups do not raise concurrency.
Luma's Build tier
Luma's page gives a small table for the Build tier and says higher limits are available through its Scale Plan form. It does not describe 429 responses or rate-limit headers on that page, so build your client to read the actual response.
| Item | Luma Build tier | Sume |
|---|---|---|
| Video concurrency | 10 (Ray) | Plan-only concurrency_limit; top-ups do not raise it |
| Request rate | 20 per minute (Ray) | 429 rate_limited with ratelimit-* and retry-after when present |
| Monthly cap | $5000 usage limit | Workspace USD balance; no monthly cap named on these pages |
| Extra capacity | Scale Plan form | Plan change or an admin override |
| When full | Not described | Accepted as queued, then 429 queue_full at capacity |
How Sume queues instead
Sume's admission page treats concurrency as a dispatch limit, not a submit limit. A valid job is accepted as queued and moves to processing when a slot opens. Queue capacity defaults to max(3, concurrency_limit x 5), and generation_limits on the API shows your effective values, including queue_capacity_remaining and a wave_size_hint for sizing a batch.
Sizing a batch
Size your batch from the remaining queue capacity, not from a concurrency guess. The admission page warns that wave_size_hint is a submission-wave hint only and never a processing width, so do not use it to size in-flight work.
- Read your limits before a large run, not after the first
429. - Submit in waves that fit the remaining queue capacity.
- Back off on
rate_limitedusingretry-after. - Use an idempotency key on every submit that may be retried.
Where you will feel each limit
A request-per-minute limit like Luma's mostly shapes how fast you can submit. On Sume the submit side is looser and the limit you feel is the processing concurrency of your plan.
Sources
Related posts
More in Developers
- Luma Dream Machine API prompt rules: 3 to 5000 characters
Luma's Dream Machine API rejects prompts under 3 or over 5000 characters, and loop with keyframes. The pre-submit errors, and where Sume's checks live.
- Luma generation states and callback_url vs Sume statuses
Luma's Dream Machine API reports dreaming, completed and failed, with a callback_url POST. How that maps to Sume's queued, processing and completed.
- Luma modify video: ray-flash-2 15 seconds, ray-2 10, 100 MB
Luma's Modify Video allows 15 seconds on ray-flash-2 and 10 on ray-2, with a 100 MB source. Sume edits video with video_url on gemini-omni-flash-1.1.
- Luma failure_reason moderation messages vs Sume generation_rejected
Luma reports moderation and dispatch failures as strings in failure_reason. Sume returns a job error category and next action. A map for handling both.
Written by Sume