OpenAI 2,000 batches an hour vs Sume's write budget per minute

OpenAI's Batch API allows 2,000 batch creations per hour. Sume budgets requests per minute by plan: 120 writes on Free to 1,200 on Scale.

5 min readSume
All posts

OpenAI's Batch API lets you create at most 2,000 batches per hour, plus model-specific limits on queued tokens set in your account settings. Sume instead gives every key a per-minute request budget by plan, split into reads and writes: 120 writes and 4,800 reads on Free, up to 1,200 writes and 48,000 reads on Scale. Creating a bulk queue spends one write, whatever the item count.

OpenAI's limits are from its Batch API guide; Sume's are from Errors and spend.

What are Sume's per-minute budgets?

A read is any GET: the receipt, status_url, events_url, result_url, and the Format and run lists. A write is everything else, such as creating runs and queues, cancel and redeliver. Reads get forty times the write number, so polling cannot starve your creates.

Sume request budgets per minute by plan (read 2026-10-02)
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000
EnterpriseContracted; Scale until provisionedContracted

How do the two limits compare in practice?

They measure different things. OpenAI counts batches over an hour; Sume counts requests over a minute and says nothing about hourly caps in the text we read. A 100-item queue is one write on Sume, so the request budget is rarely what slows a bulk submit. The docs say request rate is not generation capacity: how many generations run at once is governed by the plan's concurrency limit, and raising your request rate does not raise it.

So on Sume the thing that bounds a batch is the queue's concurrency window (1 to 16) and workspace generation concurrency, not the writes per minute.

What do I read on a 429?

A 429 rate_limited names its budget in error.details.scope, read or write, and sends retry-after. Every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and the docs say to pace on those headers rather than counting requests yourself. A 429 or 503 inside a poll loop is transient: abandoning the loop does not stop the run or its spend, so back off and poll again.

import os, time, requests

def get(url: str) -> dict:
    for attempt in range(6):
        r = requests.get(url, timeout=30, headers={
            "Authorization": f"Bearer {os.environ['SUME_API_KEY']}"})
        if r.status_code == 429:
            wait = float(r.headers.get("retry-after", 2 ** attempt))
            time.sleep(wait)
            continue
        r.raise_for_status()
        return r.json()
    raise RuntimeError("still rate limited")

print(get(os.environ["QUEUE_URL"])["data"]["counts"])

How should I pace a large submit?

Poll queues with backoff rather than a fixed one-second sleep: each child run is minutes of work when the Format makes video, so a fast loop buys nothing and spends read budget. Do not retry a 403 insufficient_scope in a loop; the docs call it the most common and most expensive mistake, and it needs a new key rather than a wait.

If you outgrow your plan's writes per minute, the table above is the ceiling; Enterprise limits are contracted. See pacing bulk submits against generation limits.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume