service_account_policy_unavailable 503: the message says temporary
The message says temporarily unavailable, but the envelope marks this 503 retryable false with contact_support. Retry cautiously and alert on repeats.

service_account_policy_unavailable is a 503 with the message "Service account policy enforcement is temporarily unavailable." The word temporarily suggests a retry, but the public envelope derived from the status says otherwise: category: runtime_unavailable, retryable: false, next_action: contact_support. The two disagree, so decide with a bounded retry and an alert, not an endless loop.
When the API raises it
On a paid submit with a service-account key, the API first checks the key, model, operation and header rules, and then needs a policy store to count spend and queued jobs. If that store is not available, it cannot tell whether a cap would be crossed, so it refuses the request instead of letting it through unmetered. That is a fail-closed design: nothing is reserved and nothing is billed.
Why the flag says false
The actionability mapping has a rule for 501 and 503 codes without a specific entry: they become runtime_unavailable and contact_support. A few 503s are carved out as retryable, such as deploy_draining, database_busy and capacity exhaustion, because those have known recovery. This code is not on that list.
| Code | retryable | next_action |
|---|---|---|
| service_account_policy_unavailable | false | contact_support |
| spend_approval_store_misconfigured | false | contact_support |
| deploy_draining | true | retry_later |
| database_busy | true | retry_later |
A bounded retry
Honor the flag, and add a small number of delayed attempts because the message says the condition is temporary. Then escalate. This sketch runs offline:
import json
BODY = '{"error": {"code": "service_account_policy_unavailable", "retryable": false, "next_action": "contact_support"}}'
def plan(err: dict, attempt: int) -> str:
if err["retryable"]:
return "retry with backoff"
if err["code"].endswith("_unavailable") and attempt < 2:
return f"one cautious retry in {30 * (attempt + 1)}s"
return "alert on-call with request_id"
for n in range(3):
print(n, plan(json.loads(BODY)["error"], n))Monitoring
Count this code separately from ordinary 5xx. A single occurrence is noise, and a run of them is an incident, because every paid submit on service-account keys is being refused. Alert on the rate of this code, and include the request_id of the latest one in the page. Because nothing is billed on a refusal, the main cost of an outage is delay, so a visible queue length on your side is the best signal for your own users.
Keep it idempotent
Use the same Idempotency-Key for each attempt: no job exists, so the same key cannot duplicate work, and the retry stays safe. If it fails twice, hold the work in your own queue and report the request_id. The Errors and credits page shows the envelope fields.
If you own the integration end to end, document this code in your runbook next to the cap errors, with the expected action for each. The runbook entry should say: bounded retry, same idempotency key, alert on repeats and attach the request id. A short runbook line saves minutes during an incident, when no one wants to read source.
Sources
Related posts
More in Developers
- service_account_rate_limited 429: read the minute and hour headers
A service-account key has an optional minute window and hour window. The 429 names the window, and x-sume-service-account headers show what is left.
- Service-account 402: daily, monthly or per-end-user spend cap hit
Three different 402 codes mean three different caps. Read details.cap_usd_micros and current_usd_micros to see which window blocked the request.
- language_code or auto-detect on Sume STT? A two-arm test on your clips
AssemblyAI reports 8.4% mean WER over 18 FLEURS languages. For your audio, run each clip twice on Sume STT, with and without language_code, and compare.
- SHA-256 manifest for a batch of 30-second clips: spot bad downloads
After downloading many 30-second Sume clips, write a manifest of size and SHA-256 per job id so a rerun skips good files and re-fetches bad ones. Python.
Written by Sume