Sume API 503 deploy_draining: retry after 5 seconds in Python
A redeploy answers 503 deploy_draining with retry-after. Why it is safe to retry, why a POST still needs an Idempotency-Key, and a stdlib Python retry loop.

When the Sume API redeploys, the replica that is shutting down answers every new request with 503 and the code deploy_draining. The response carries a retry-after header and the error envelope says retryable: true with retry_after_seconds (5 seconds unless the server sets another value). The retry lands on the replacement replica, so waiting a few seconds and sending the same request again is the correct response.
Why this 503 is different from the other 503s
Most 503 codes on the API mean a dependency is missing, and the envelope marks them retryable: false with a contact-support action. deploy_draining is the exception, and the API classifies it separately so that a client reading retryable does not give up. The drain guard runs before rate limiting and auth, so a refused request spends neither budget.
A POST can be a repeat
The same 503 is sent to requests that were still running when the drain cutoff arrived. For a submit, that means the job may already exist when you see the error. Send an Idempotency-Key (1 to 255 characters) on every POST that you might retry, and reuse the same key on each attempt. A retry with the same key and the same body returns the original job instead of creating a second one.
Retry loop
This helper uses only the standard library. It reads the wait from the envelope, falls back to the header, and retries only deploy_draining. Any other error is raised for the caller to classify.
import json, os, time, urllib.error, urllib.request, uuid
BASE = os.environ.get("SUME_BASE_URL", "https://api.sume.com")
def call(method, path, body=None, attempts=4):
data = json.dumps(body).encode() if body is not None else None
headers = {"x-api-key": os.environ["SUME_API_KEY"]}
if data is not None:
headers["content-type"] = "application/json"
headers["idempotency-key"] = str(uuid.uuid4()) # one key for every attempt
for attempt in range(attempts):
req = urllib.request.Request(BASE + path, data=data, method=method, headers=headers)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
return json.load(resp)
except urllib.error.HTTPError as exc:
err = json.load(exc).get("error", {})
draining = exc.code == 503 and err.get("code") == "deploy_draining"
if not draining or attempt == attempts - 1:
raise
wait = err.get("retry_after_seconds") or int(exc.headers.get("retry-after", 5))
time.sleep(wait)What to log
Log error.request_id on every failed attempt. If four attempts in a row return deploy_draining, which is about 15 seconds of waiting with the default delay, stop and surface the error. A drain is normally over well before that. For a long-running generation, the job state lives on the server, so poll its status_url instead of resubmitting.
Sources
Related posts
More in Developers
- A blank Idempotency-Key on Sume is ignored, not rejected: guard it
An empty or whitespace Idempotency-Key header is treated as no key at all, so a retry can create a second job. Build a key that cannot be blank in Python.
- Parse the Sume error envelope into a Python dataclass and exception
Sume errors share one envelope: code, request_id, retryable, retry_after_seconds, next_action, category and stage. Turn it into a typed Python exception.
- Get the Sume webhook secret with GET /v1/webhooks/signing-secret
Read the workspace webhook secret with an account:read key and load it as SUME_COM_WEBHOOK_SIGNING_SECRET without printing it. A Python deploy step.
- GET /v1/jobs 400 unknown_parameter: a typo'd filter no longer widens
A misspelled query key on GET /v1/jobs, such as state for status, now returns 400 with a suggestion instead of a full unfiltered page. Handle it in TypeScript.
Written by Sume