Sume API 503 codes: which ones are retryable and which are not
Four 503 responses look alike but differ in the retryable flag: deploy_draining, database_busy, provider_capacity_exceeded and any _not_configured code.

A bare 503 tells you very little. The Sume API puts the real answer in the error envelope: error.code, error.retryable, error.retry_after_seconds and error.next_action. The status code alone would send you to a retry loop for some failures that retrying cannot fix. Read the envelope, not the status.
The four cases
The API classifies these 503 responses in its error handler. The retry delays below are the values the server uses when it does not set one itself.
| error.code | retryable | retry_after_seconds | Meaning |
|---|---|---|---|
| deploy_draining | true | 5 | A replica is shutting down for a redeploy. |
| database_busy | true | 1 | The API's database admission queue refused the request. |
| provider_capacity_exceeded | true | 30 | The provider has no capacity for the submit. |
| any code ending in _not_configured | false | null | A runtime dependency is missing. Retrying does not help. |
Rule of thumb
deploy_draininganddatabase_busyare the API's own short blips. Wait the stated seconds and resend with the same Idempotency-Key.provider_capacity_exceededis the same family as the 429queue_full. Wait 30 seconds, or pace the wave using thegeneration_limitsobject.- A 503 that is not one of those three is marked
retryable: falsewith a contact-support action. Keep therequest_idand write to support instead of looping.
A classifier
The function below turns an error envelope into a decision. It never inspects the HTTP status, so it keeps working if a code moves to another status.
def decide(envelope):
"""Return ('retry', seconds) or ('stop', reason) for a Sume error envelope."""
err = envelope.get("error", {})
if err.get("retryable") is True:
return ("retry", err.get("retry_after_seconds") or 5)
return ("stop", err.get("next_action") or err.get("code") or "unknown")
drain = {"error": {"code": "deploy_draining", "retryable": True, "retry_after_seconds": 5}}
missing = {"error": {"code": "provider_not_configured", "retryable": False,
"retry_after_seconds": None, "next_action": "contact_support"}}
print(decide(drain)) # ('retry', 5)
print(decide(missing)) # ('stop', 'contact_support')What the docs say
The public error table lists 503 with provider_not_configured, provider_capacity_exceeded or a storage configuration error, and describes it as a runtime dependency that is unavailable or at capacity. The retryable flag in the envelope is what separates the two groups, so treat it as the contract.
Sources
Related posts
More in Developers
- Sume API 503 deploy_draining: retry after 5 seconds in Python
A redeploy answers 503 deploy_draining with retry-after. Why it is safe to retry, why a POST still needs an Idempotency-Key, and a stdlib Python retry loop.
- A blank Idempotency-Key on Sume is ignored, not rejected: guard it
An empty or whitespace Idempotency-Key header is treated as no key at all, so a retry can create a second job. Build a key that cannot be blank in Python.
- Parse the Sume error envelope into a Python dataclass and exception
Sume errors share one envelope: code, request_id, retryable, retry_after_seconds, next_action, category and stage. Turn it into a typed Python exception.
- Get the Sume webhook secret with GET /v1/webhooks/signing-secret
Read the workspace webhook secret with an account:read key and load it as SUME_COM_WEBHOOK_SIGNING_SECRET without printing it. A Python deploy step.
Written by Sume