How many Format runs can one API key poll at once?

Read budgets per minute by Sume plan, turned into a count of Format runs you can poll at 5-second and 30-second intervals, and why a bulk queue saves writes.

5 min readSume
All posts

On Sume's Free plan one API key can read 4800 times a minute, so it can poll about 400 Format runs every 5 seconds or 2400 runs every 30 seconds. Pro, Startup and Scale give 2.5, 5 and 10 times that. Creating runs is the tighter limit: Free allows 120 writes a minute, so 100 single creates already use most of it.

The numbers come from the rate-limit table on Authentication and rate limits. The division below is mine, and it assumes one GET per run per poll and no other read traffic on the key.

The budgets and the division

A read is any GET or HEAD, including a poll of a run's status_url. Run creation, cancellation and uploads are writes. The docs say the plan number is the write number, and reads get forty times that in a separate bucket. Enterprise uses the Scale row until Sume provisions a contracted number.

Runs a key can poll if each poll is one read (limits from the docs, read 2026-10-10; division is arithmetic)
PlanWrites / minReads / minRuns polled every 5 s (12 reads/min)Runs polled every 30 s (2 reads/min)
Free12048004002400
Pro3001200010006000
Startup60024000200012000
Scale120048000400024000

Writes are the real ceiling

A Format run lives up to 90 minutes before expires_at, so a batch of single creates holds its reads for a long time, but the creates themselves are limited by the smaller write bucket. One hundred POST …/runs calls use 100 of the Free plan's 120 writes in that minute. The same hundred items in one bulk queue use one write, because the server creates the children.

A queue accepts 1 to 100 items and a concurrency window of 1 to 16. You then poll one queue resource and read children only when you need their receipts. Bulk runs also have no queue-level webhook, so polling is the intended path there.

What the limit does not cap

The docs separate request rate from generation capacity. Your plan's concurrency limit controls how many generations run at the same time, and raising your request rate does not raise it. A key that polls ten thousand runs a minute still waits on the same generation concurrency.

Do not count requests by hand. Read ratelimit-remaining, and on a 429 wait for retry-after. The 429 body names the exhausted bucket in error.details.scope, either read or write, so you can tell a polling storm from a creation burst.

Cheaper ways to watch runs

Prefer a webhook for single runs, since a format.run.terminal delivery replaces the polling loop entirely. For bulk queues, poll the queue on a slow interval and widen it as the queue drains.

If you must poll runs, back off by run age: a clip-heavy run rarely finishes in the first minute, so a 5-second loop wastes budget that a 30-second loop would keep.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume