OpenAI 2,000 batches an hour vs Sume's write budget per minute
OpenAI's Batch API allows 2,000 batch creations per hour. Sume budgets requests per minute by plan: 120 writes on Free to 1,200 on Scale.

OpenAI's Batch API lets you create at most 2,000 batches per hour, plus model-specific limits on queued tokens set in your account settings. Sume instead gives every key a per-minute request budget by plan, split into reads and writes: 120 writes and 4,800 reads on Free, up to 1,200 writes and 48,000 reads on Scale. Creating a bulk queue spends one write, whatever the item count.
OpenAI's limits are from its Batch API guide; Sume's are from Errors and spend.
What are Sume's per-minute budgets?
A read is any GET: the receipt, status_url, events_url, result_url, and the Format and run lists. A write is everything else, such as creating runs and queues, cancel and redeliver. Reads get forty times the write number, so polling cannot starve your creates.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Enterprise | Contracted; Scale until provisioned | Contracted |
How do the two limits compare in practice?
They measure different things. OpenAI counts batches over an hour; Sume counts requests over a minute and says nothing about hourly caps in the text we read. A 100-item queue is one write on Sume, so the request budget is rarely what slows a bulk submit. The docs say request rate is not generation capacity: how many generations run at once is governed by the plan's concurrency limit, and raising your request rate does not raise it.
So on Sume the thing that bounds a batch is the queue's concurrency window (1 to 16) and workspace generation concurrency, not the writes per minute.
What do I read on a 429?
A 429 rate_limited names its budget in error.details.scope, read or write, and sends retry-after. Every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and the docs say to pace on those headers rather than counting requests yourself. A 429 or 503 inside a poll loop is transient: abandoning the loop does not stop the run or its spend, so back off and poll again.
import os, time, requests
def get(url: str) -> dict:
for attempt in range(6):
r = requests.get(url, timeout=30, headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"})
if r.status_code == 429:
wait = float(r.headers.get("retry-after", 2 ** attempt))
time.sleep(wait)
continue
r.raise_for_status()
return r.json()
raise RuntimeError("still rate limited")
print(get(os.environ["QUEUE_URL"])["data"]["counts"])How should I pace a large submit?
Poll queues with backoff rather than a fixed one-second sleep: each child run is minutes of work when the Format makes video, so a fast loop buys nothing and spends read budget. Do not retry a 403 insufficient_scope in a loop; the docs call it the most common and most expensive mistake, and it needs a new key rather than a wait.
If you outgrow your plan's writes per minute, the table above is the ceiling; Enterprise limits are contracted. See pacing bulk submits against generation limits.
Sources
Related posts
More in Comparisons
- OpenRouter models fallback array and 3-entry limit vs Sume
OpenRouter's models array tries the next model on downtime, rate limits or moderation; fallbacks allows 3. Sume's allow_fallbacks has no effect.
- OpenRouter provider.sort and max_price vs Sume's inert sort
OpenRouter's provider.sort picks price, throughput or latency and turns off load balancing. On Sume's image route, sort is accepted and changes nothing.
- Pexels free stock footage vs generated B-roll: what to use when
Compare Pexels licensed footage with generated B-roll: licence limits, control, cost and where each wins, with Pexels terms to re-check.
- Pictory video minutes per dollar vs Sume timeline render
Pictory lists 200 to 1,800 video minutes a month at $0.066 to $0.145 per minute. Sume renders a timeline at $0.10 per output minute. Dated 2026-10-01.
Written by Sume