Luma Build tier: 10 concurrent Ray jobs, 20 requests a minute

Luma's Build tier allows 10 concurrent Ray video jobs, 20 requests a minute and $5000 a month. How Sume's plan concurrency and queue capacity differ.

5 min readSume
All posts

On Luma's documented Build tier, the Ray video model allows 10 concurrent generations and 20 requests per minute, the Photon image models allow 40 and 80, and the tier has a usage limit of $5000 per month. Sume describes limits differently: generation concurrency is set by your plan, queue capacity sits on top of it, and prepaid top-ups do not raise concurrency.

Luma's Build tier

Luma's page gives a small table for the Build tier and says higher limits are available through its Scale Plan form. It does not describe 429 responses or rate-limit headers on that page, so build your client to read the actual response.

Rate limits, read 2026-10-02
ItemLuma Build tierSume
Video concurrency10 (Ray)Plan-only concurrency_limit; top-ups do not raise it
Request rate20 per minute (Ray)429 rate_limited with ratelimit-* and retry-after when present
Monthly cap$5000 usage limitWorkspace USD balance; no monthly cap named on these pages
Extra capacityScale Plan formPlan change or an admin override
When fullNot describedAccepted as queued, then 429 queue_full at capacity

How Sume queues instead

Sume's admission page treats concurrency as a dispatch limit, not a submit limit. A valid job is accepted as queued and moves to processing when a slot opens. Queue capacity defaults to max(3, concurrency_limit x 5), and generation_limits on the API shows your effective values, including queue_capacity_remaining and a wave_size_hint for sizing a batch.

Sizing a batch

Size your batch from the remaining queue capacity, not from a concurrency guess. The admission page warns that wave_size_hint is a submission-wave hint only and never a processing width, so do not use it to size in-flight work.

  • Read your limits before a large run, not after the first 429.
  • Submit in waves that fit the remaining queue capacity.
  • Back off on rate_limited using retry-after.
  • Use an idempotency key on every submit that may be retried.

Where you will feel each limit

A request-per-minute limit like Luma's mostly shapes how fast you can submit. On Sume the submit side is looser and the limit you feel is the processing concurrency of your plan.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume