How many Format runs per minute can my Sume plan start?
Sume limits writes per minute by plan, from 120 on Free to 1200 on Scale, with reads at 40 times that. Read the rate-limit headers and back off on 429.

Your Sume plan sets a write budget per minute: 120 on Free, 300 on Pro, 600 on Startup, and 1200 on Scale. Creating a Format run is a write, and reads get 40 times the write budget, so the create rate is the limit you will meet first.
Four response headers tell you where you stand: ratelimit-limit, ratelimit-remaining, ratelimit-reset, and retry-after on a 429.
These budgets are per key plan, not per Format, so every integration sharing a key shares the minute. Give each service its own key where you can, so one noisy job cannot starve another.
The numbers
The Formats docs give writes and reads per minute for each plan. Reads are 40 times writes, so the read column is 4800, 12000, 24000, and 48000. A bulk queue creates up to 100 runs from one request, but the items still start at the concurrency you set, from 1 to 16.
A quick check of the arithmetic: Pro at 300 writes a minute is 18,000 an hour, and Scale at 1200 is 72,000. A daily batch of a few thousand runs fits any plan if you spread it, and breaks Free if you do not.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
Polling cost
Polling spends reads. A 30 minute run polled with a delay doubling up to 60 seconds takes about 35 status reads, so a Free plan's 4800 reads a minute supports a very large number of runs polled at once. In practice your creates and your cancels, which are writes, are the budget to watch.
If you can use webhooks, you will spend far fewer reads. Use a poll only as a safety net.
Handling a 429
When a call returns 429, wait for retry-after seconds before retrying, and spread the retry with jitter so many workers do not wake together. Use the same Idempotency-Key on the retry of a create so a request that did land is replayed, not duplicated. The Call a Format page lists the headers.
A bulk queue is the way to smooth a burst. Submit up to 100 items in one create, set concurrency to what your plan and providers can hold, and let the queue pace the runs. See Bulk runs.
Do not retry in a tight loop. A client that ignores retry-after and tries again immediately makes the next window worse, and a few of them together look like an attack.
Capacity planning
Estimate your peak create rate, not the average. A nightly job that creates 500 runs at once will exceed 120 writes a minute on Free in a few seconds, so batch it or pick a plan to fit. Track ratelimit-remaining in your logs, and alert when it falls below a tenth of the limit during normal traffic.
Review the budget each quarter. Usage grows quietly, and moving to a bigger plan is cheaper than the engineering time spent on tuning around a limit that no longer fits.
Sources
Related posts
More in Developers
- HyperFrames check via the Sume API: caption collisions pre-render
Send check with caption_zone to POST /v1/hyperframes-previews and get findings, contrast and overlap reports. A failing check is still a completed job.
- image_not_fetchable on a Sume image edit: reference URL checklist
A Sume image edit failed with image_not_fetchable or input_media_unreachable. What the docs say the error means and a checklist for the reference URL.
- Image API returned 202, not an image: one Python handler for both
POST /v1/images waits 30 seconds, then returns a 202 job envelope. A Python handler that reads the status code, polls the job and returns image URLs either way.
- ky retry on POST for Sume: Idempotency-Key, 40 s timeout, v2 baseUrl
ky does not retry POST by default and times out at 10 s, but Sume sync can hold 30 s. A tested ky v2 config with a stable Idempotency-Key and no 429 retries.
Written by Sume