How many Format runs can one API key poll at once?
Read budgets per minute by Sume plan, turned into a count of Format runs you can poll at 5-second and 30-second intervals, and why a bulk queue saves writes.

On Sume's Free plan one API key can read 4800 times a minute, so it can poll about 400 Format runs every 5 seconds or 2400 runs every 30 seconds. Pro, Startup and Scale give 2.5, 5 and 10 times that. Creating runs is the tighter limit: Free allows 120 writes a minute, so 100 single creates already use most of it.
The numbers come from the rate-limit table on Authentication and rate limits. The division below is mine, and it assumes one GET per run per poll and no other read traffic on the key.
The budgets and the division
A read is any GET or HEAD, including a poll of a run's status_url. Run creation, cancellation and uploads are writes. The docs say the plan number is the write number, and reads get forty times that in a separate bucket. Enterprise uses the Scale row until Sume provisions a contracted number.
| Plan | Writes / min | Reads / min | Runs polled every 5 s (12 reads/min) | Runs polled every 30 s (2 reads/min) |
|---|---|---|---|---|
| Free | 120 | 4800 | 400 | 2400 |
| Pro | 300 | 12000 | 1000 | 6000 |
| Startup | 600 | 24000 | 2000 | 12000 |
| Scale | 1200 | 48000 | 4000 | 24000 |
Writes are the real ceiling
A Format run lives up to 90 minutes before expires_at, so a batch of single creates holds its reads for a long time, but the creates themselves are limited by the smaller write bucket. One hundred POST …/runs calls use 100 of the Free plan's 120 writes in that minute. The same hundred items in one bulk queue use one write, because the server creates the children.
A queue accepts 1 to 100 items and a concurrency window of 1 to 16. You then poll one queue resource and read children only when you need their receipts. Bulk runs also have no queue-level webhook, so polling is the intended path there.
What the limit does not cap
The docs separate request rate from generation capacity. Your plan's concurrency limit controls how many generations run at the same time, and raising your request rate does not raise it. A key that polls ten thousand runs a minute still waits on the same generation concurrency.
Do not count requests by hand. Read ratelimit-remaining, and on a 429 wait for retry-after. The 429 body names the exhausted bucket in error.details.scope, either read or write, so you can tell a polling storm from a creation burst.
Cheaper ways to watch runs
Prefer a webhook for single runs, since a format.run.terminal delivery replaces the polling loop entirely. For bulk queues, poll the queue on a slow interval and widen it as the queue drains.
If you must poll runs, back off by run age: a clip-heavy run rarely finishes in the first minute, so a 5-second loop wastes budget that a 30-second loop would keep.
Sources
Related posts
More in Formats
- Is there an endpoint to list all Format runs? No, build an index
Sume lists runs per Format with a cursor, but there is no GET /v1/format-runs across Formats. Here is the small run index that fills the gap, and what to store.
- Make a partial Format result legal in your output_schema
Sume Format output_schema has no optional properties. Use nullable unions, SumeMediaFile refs and honest nulls so a run that makes 2 of 3 clips still returns.
- Re-run a Format on purpose: version the Idempotency-Key
How Sume Format idempotency works, why a fresh uuid per request defeats it, and how an order id plus a version number gives safe retries and deliberate re-runs.
- Sume Format instruction limit: why long briefs belong in input
A Format instruction accepts 8000 characters but only about the first 4000 reach the prompt. Put long briefs in the input field, which is stored whole as data.
Written by Sume