service_account_rate_limited 429: read the minute and hour headers

A service-account key has an optional minute window and hour window. The 429 names the window, and x-sume-service-account headers show what is left.

4 min readSume
All posts

service_account_rate_limited is a request-rate 429 on a service-account key. Unlike the job-count caps, it has a Retry-After header, and details names which window you exhausted, minute or hour. The limits are set per key, and a window exists only if it has a limit, so a key can have one, both or neither.

Headers on every checked response

Whenever a window is checked, the response carries its limit and what is left, even when the request succeeds. That lets a client slow down before it hits the wall.

Service-account rate-limit headers (Sume API source, read 2026-10-05)
HeaderMeaning
x-sume-service-account-minute-limitRequests allowed in the minute window
x-sume-service-account-minute-remainingRequests left in that window
x-sume-service-account-hour-limitRequests allowed in the hour window
x-sume-service-account-hour-remainingRequests left in the hour window
retry-afterSeconds to wait, set on the 429 only

The 429 body

details has window, limit and retry_after_seconds. If the rate-limit store was degraded, the check still returns an answer and adds rate_limit_degraded: true, which tells you the count may be approximate and is not a reason to hammer. Because the minute window is checked before the hour window, a minute failure comes back first even when the hour budget is also short.

Honor the header, with a floor

Read retry-after first, then the body value, and add a small floor so a 0 does not turn into a tight loop:

import json, time

HEADERS = {"retry-after": "12"}
BODY = '{"error": {"code": "service_account_rate_limited", "details": {"window": "minute", "limit": 60, "retry_after_seconds": 12}}}'

err = json.loads(BODY)["error"]
wait = int(HEADERS.get("retry-after") or err["details"]["retry_after_seconds"] or 1)
wait = max(wait, 1)
print(f"{err['details']['window']} window full, sleeping {wait}s")
time.sleep(min(wait, 2))  # demo cap; use the full wait in real code

Pacing before you hit it

Because the remaining counters ride on normal responses, a client can pace itself with no extra calls. A simple rule is to hold new requests when x-sume-service-account-minute-remaining reaches 0 and release them at the next window boundary, and to treat the hour counter as a budget you spread, not a burst allowance. Keep one limiter per key, not per process, or each process will believe it has the full budget.

If many callers share a key, consider a queue in front of the Sume client so that bursts are smoothed by you and not rejected by the key's policy. A rejected request still counts against your own retry budget, and a polite client avoids it.

Alerting

Alert on the hour window, not the minute window. A minute refusal clears itself within the minute, but an hour refusal can hold you for a long time, and it signals that your steady-state rate is too high for the key. Chart the remaining counters, and the shape of the problem is visible well before the first 429.

Not the same as the workspace limiter

The workspace-level rate_limited is a separate limiter with its own headers. If your headers start with x-sume-service-account-, the key's own policy fired, and changing your workspace plan will not move it. See Errors and credits for the general retry rules.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume