service_account_rate_limited 429: read the minute and hour headers
A service-account key has an optional minute window and hour window. The 429 names the window, and x-sume-service-account headers show what is left.

service_account_rate_limited is a request-rate 429 on a service-account key. Unlike the job-count caps, it has a Retry-After header, and details names which window you exhausted, minute or hour. The limits are set per key, and a window exists only if it has a limit, so a key can have one, both or neither.
Headers on every checked response
Whenever a window is checked, the response carries its limit and what is left, even when the request succeeds. That lets a client slow down before it hits the wall.
| Header | Meaning |
|---|---|
| x-sume-service-account-minute-limit | Requests allowed in the minute window |
| x-sume-service-account-minute-remaining | Requests left in that window |
| x-sume-service-account-hour-limit | Requests allowed in the hour window |
| x-sume-service-account-hour-remaining | Requests left in the hour window |
| retry-after | Seconds to wait, set on the 429 only |
The 429 body
details has window, limit and retry_after_seconds. If the rate-limit store was degraded, the check still returns an answer and adds rate_limit_degraded: true, which tells you the count may be approximate and is not a reason to hammer. Because the minute window is checked before the hour window, a minute failure comes back first even when the hour budget is also short.
Honor the header, with a floor
Read retry-after first, then the body value, and add a small floor so a 0 does not turn into a tight loop:
import json, time
HEADERS = {"retry-after": "12"}
BODY = '{"error": {"code": "service_account_rate_limited", "details": {"window": "minute", "limit": 60, "retry_after_seconds": 12}}}'
err = json.loads(BODY)["error"]
wait = int(HEADERS.get("retry-after") or err["details"]["retry_after_seconds"] or 1)
wait = max(wait, 1)
print(f"{err['details']['window']} window full, sleeping {wait}s")
time.sleep(min(wait, 2)) # demo cap; use the full wait in real codePacing before you hit it
Because the remaining counters ride on normal responses, a client can pace itself with no extra calls. A simple rule is to hold new requests when x-sume-service-account-minute-remaining reaches 0 and release them at the next window boundary, and to treat the hour counter as a budget you spread, not a burst allowance. Keep one limiter per key, not per process, or each process will believe it has the full budget.
If many callers share a key, consider a queue in front of the Sume client so that bursts are smoothed by you and not rejected by the key's policy. A rejected request still counts against your own retry budget, and a polite client avoids it.
Alerting
Alert on the hour window, not the minute window. A minute refusal clears itself within the minute, but an hour refusal can hold you for a long time, and it signals that your steady-state rate is too high for the key. Chart the remaining counters, and the shape of the problem is visible well before the first 429.
Not the same as the workspace limiter
The workspace-level rate_limited is a separate limiter with its own headers. If your headers start with x-sume-service-account-, the key's own policy fired, and changing your workspace plan will not move it. See Errors and credits for the general retry rules.
Sources
Related posts
More in Developers
- Service-account 402: daily, monthly or per-end-user spend cap hit
Three different 402 codes mean three different caps. Read details.cap_usd_micros and current_usd_micros to see which window blocked the request.
- language_code or auto-detect on Sume STT? A two-arm test on your clips
AssemblyAI reports 8.4% mean WER over 18 FLEURS languages. For your audio, run each clip twice on Sume STT, with and without language_code, and compare.
- SHA-256 manifest for a batch of 30-second clips: spot bad downloads
After downloading many 30-second Sume clips, write a manifest of size and SHA-256 per job id so a rerun skips good files and re-fetches bad ones. Python.
- Shorts series episodes number by publish date: a Python order check
YouTube numbers Shorts series episodes by publish date, so upload order is episode order. Check durations and publish times in Python before you schedule.
Written by Sume