Read Sume rate limit headers with curl: remaining, reset, retry-after
Run curl -D - on any /v1 route to read ratelimit-limit, ratelimit-remaining, ratelimit-reset, and retry-after on a 429. Anonymous reads are 4x Free writes.

Add -D - and -o /dev/null to a curl call, and the response headers print without the body. Look for ratelimit-limit, ratelimit-remaining and ratelimit-reset on every response, and for retry-after on a 429. You can try it with no key, because GET /v1/health is a public route. Authenticated calls show the budget of your plan instead of the anonymous one.
The headers describe the budget that the current request spent from. That detail matters, because a read and a write draw from different buckets. A GET for the balance shows the read bucket, and a POST that creates work shows the write bucket, so check the headers on the type of call that you worry about.
Keep this habit for any new integration. Print the headers once from the machine that will run the job, because a shared egress address, a proxy or a different plan can change what you see.
What each header says
The four headers have fixed meanings in the authentication docs. ratelimit-limit is the number of requests permitted in the current window. ratelimit-remaining is what is left. ratelimit-reset is the number of seconds until the window resets. retry-after is the number of seconds to wait, and it is sent on a 429.
Do not count requests yourself. The docs give this advice directly, because your count and the server count drift apart as soon as there is a second process, a retry or a second key. The header is the source of truth for the deployment that you call, and the read multiple is a deployment setting, so a preview deployment can differ from the numbers in the docs.
What to expect by plan
Rate limits are per key per minute, with a write bucket and a separate read bucket. The read bucket is 40 times the write bucket for keyed accounts.
Treat the numbers as a floor for planning, not as a promise. The read multiple is a deployment setting, so a self-hosted or preview deployment can use other values. The ratelimit-limit header on the response is always the authority for the deployment you call, and a script that adapts to the header keeps working when a limit changes.
One more point for agents and automations. An MCP tool call spends the write budget once, for the run that it creates, and it does not spend the write budget for the JSON-RPC request that carried it, and a status poll over MCP spends no write budget at all.
| Plan | Write bucket | Read bucket |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Unauthenticated, per client IP | Free write rate | 4 times the write rate |
Two curl commands
Run the two commands below. The first reads the public health route with no key and prints only the headers that matter. The second does the same for the balance route with your key, and it needs the key in an environment variable.
The anonymous example needs no key, so you can run it anywhere. Unauthenticated calls are limited per client IP at the Free rate, and their read bucket is four times the write rate, not forty, so the figure you see is far lower than the figure for a keyed account.
A small script can sleep until the reset when ratelimit-remaining reaches zero. That is gentler than waiting for the 429, and it keeps your own logs free of errors that were avoidable. When you do get a 429, prefer retry-after to your own guess.
# no key needed: a public route
curl -s -D - -o /dev/null https://api.sume.com/v1/health | grep -i -E '^(ratelimit|retry-after|x-sume-request-id)'
# with a key: pick ONE credential header, never both
curl -s -D - -o /dev/null https://api.sume.com/v1/balance \
-H "x-api-key: $SUME_API_KEY" | grep -i -E '^(ratelimit|retry-after|x-sume-request-id)'
When you get a 429
On a 429, read error.details.scope in the body. It is read or write, and it tells you which bucket you emptied. The code is rate_limited for request volume, and the docs say to use retry-after for the backoff when it is present. A 429 with the code queue_full is different. It means the generation queue is full, and polling more slowly will not fix it.
The x-sume-request-id header is worth capturing as well. It is exposed next to the rate limit headers, and it is the id to quote when you ask support about one call.
Sources
Related posts
More in Developers
- H3 Max Recast and the 15-second shot rule: split a long take first
Recast accepts 5 to 30 seconds of source, but no single shot longer than 15 seconds. Cut a longer take with Sume's Video Trim, then recast each part.
- Missed Sume video webhook? Poll first, redeliver second
A missed callback does not mean a lost job. A Python sweeper polls open ids, then calls POST /v1/jobs/{id}/webhook/redeliver for the terminal ones.
- Recraft erase, outpaint and inpaint calls vs one edit route on Sume
Recraft has separate inpaint, outpaint and erase calls. Sume has one /v1/images route with references and mask_url (ChatGPT Image 2.5 only), plus RMBG.
- Redelivered job.completed must not start the trim twice
Redeliver re-sends the same job's terminal event with a fresh signature. Dedupe on job_id and derive the next step's Idempotency-Key from it.
Written by Sume