Ramp up API traffic gradually: pace Sume submits to the write budget

OpenAI now returns a slow_down 429 when traffic grows too fast. On Sume, pace paid submits to your plan's write budget and read ratelimit-remaining.

4 min readSume
All posts

Ramp up by sending paid submits at a steady rate below your plan's per-minute write budget, and let ratelimit-remaining tell you how much room is left instead of counting requests yourself. Start low, raise the rate in steps, and back off on retry-after when a 429 arrives.

The trigger is OpenAI's changelog entry of Sep 2: traffic that increases too quickly can return a 429 with the slow_down code, and when Retry-After is present you wait at least that long. Sume's limits are different in shape, so this page maps the habit onto Sume's rate limits, read 2026-09-30.

What budget does a Sume key get?

Every key gets a per-minute request budget across all of /v1, set by the workspace's plan. Reads and writes are separate budgets, so polling status cannot 429 your own submits.

Writes and reads per minute by plan, from the Sume authentication docs, read 2026-09-30.
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000
EnterpriseContact salesContact sales

Which calls spend the write budget?

Creating runs, cancelling and uploads are writes. Any GET or HEAD is a read, including polling status_url. An MCP tool call spends the write budget once for the run it creates, and a jobs_status poll over MCP spends none. A batch of paid submits is therefore bounded by the plan's write number, not by how often you poll.

How do I pace a batch of submits?

Pick a target below the write number, for example half of it for the first minute, and hold it. Read ratelimit-remaining and ratelimit-reset on each response; they describe the budget the request spent from. If a 429 comes back, its error.details.scope says read or write, and retry-after says how long to wait. Raise the rate only after a window passes without a 429.

Send an Idempotency-Key on every submit so a retried request replays instead of creating a second job; see idempotency keys for AI video APIs.

Why can a slow ramp still return 429?

Request rate is not generation capacity. How many generations run at once follows your plan's concurrency limit on the generation_limits object, and raising your request rate does not raise it. Concurrency being full is not an error by itself; it becomes a submit error only when the queue is also full, and that 429 is queue_full. The docs' client behavior is to wait for jobs to finish or cancel queued jobs, then retry with the same idempotency key. See queue full versus concurrency full.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume