Grok Imagine Lite's 10 requests per second vs Sume plan concurrency
xAI lists a 10 requests per second limit and Batch API for Grok Imagine 1.5 Lite. On Sume, a clip batch is bounded by plan concurrency and queue capacity.

On Sume, the limit that stops a video batch is not a requests-per-second number but accepted-job capacity: a Pro workspace can hold 24 paid generation jobs at once (4 processing plus 20 queued), and the 25th submit returns 429 queue_full. xAI's page for grok-imagine-video-1.5-lite lists a rate limit of 10 requests per second and Batch API support (xAI docs, read 2026-10-10).
Those are different controls, so a loop that worked against xAI's limit needs a different throttle on Sume.
The numbers on Sume
Sume's generation admission page separates four controls: generation concurrency, queue capacity, submit rate limits and balance reservation. Concurrency is a dispatch limit, not a submit limit: while queue capacity remains, valid jobs are accepted as queued. The default queue capacity is max(3, concurrency_limit x 5).
| Plan | Processing | Queue | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What a 100-clip Grok batch does
Take 100 image-to-video clips on the Grok row. On Pro, submit the first 24 and the rest must wait: the 25th submit gets 429 queue_full until a job finishes. On Startup the accepted limit is 48; on Scale it is 120, which holds the whole batch. Sume's docs say the dashboard Concurrency tab and the generation_limits.concurrency_limit field are the source of truth, not the static table.
A throttle that works
Treat queue_full as a signal to wait, not an error to retry in a tight loop. The errors page says to back off, use retry-after when it is present, and not to retry unsafe submits without an Idempotency-Key.
- Keep at most
accepted job capacityjobs in flight; submit the next when one completes. - Send an
Idempotency-Keyper clip so a retried submit returns the original job. - Use
callback_url(HTTPS) so you are not polling 100 jobs. - Check 402 before the batch: the balance must cover each reserve.
Money check
Each Grok clip reserves list times 1.25 per second at submit: a 6-second clip is 6 x $0.0125 = $0.075, so a 100-clip batch reserves $7.50 if all run at once. xAI's batch price may differ; Sume's reserve is shown in the poll usage.cost when the job ends.
Sources
Related posts
More in Developers
- How many Format runs per minute can my Sume plan start?
Sume limits writes per minute by plan, from 120 on Free to 1200 on Scale, with reads at 40 times that. Read the rate-limit headers and back off on 429.
- HyperFrames check via the Sume API: caption collisions pre-render
Send check with caption_zone to POST /v1/hyperframes-previews and get findings, contrast and overlap reports. A failing check is still a completed job.
- image_not_fetchable on a Sume image edit: reference URL checklist
A Sume image edit failed with image_not_fetchable or input_media_unreachable. What the docs say the error means and a checklist for the reference URL.
- Image API returned 202, not an image: one Python handler for both
POST /v1/images waits 30 seconds, then returns a 202 job envelope. A Python handler that reads the status code, polls the job and returns image URLs either way.
Written by Sume