Sume API 503 deploy_draining: retry after 5 seconds in Python

A redeploy answers 503 deploy_draining with retry-after. Why it is safe to retry, why a POST still needs an Idempotency-Key, and a stdlib Python retry loop.

4 min readSume
All posts

When the Sume API redeploys, the replica that is shutting down answers every new request with 503 and the code deploy_draining. The response carries a retry-after header and the error envelope says retryable: true with retry_after_seconds (5 seconds unless the server sets another value). The retry lands on the replacement replica, so waiting a few seconds and sending the same request again is the correct response.

Why this 503 is different from the other 503s

Most 503 codes on the API mean a dependency is missing, and the envelope marks them retryable: false with a contact-support action. deploy_draining is the exception, and the API classifies it separately so that a client reading retryable does not give up. The drain guard runs before rate limiting and auth, so a refused request spends neither budget.

A POST can be a repeat

The same 503 is sent to requests that were still running when the drain cutoff arrived. For a submit, that means the job may already exist when you see the error. Send an Idempotency-Key (1 to 255 characters) on every POST that you might retry, and reuse the same key on each attempt. A retry with the same key and the same body returns the original job instead of creating a second one.

Retry loop

This helper uses only the standard library. It reads the wait from the envelope, falls back to the header, and retries only deploy_draining. Any other error is raised for the caller to classify.

import json, os, time, urllib.error, urllib.request, uuid

BASE = os.environ.get("SUME_BASE_URL", "https://api.sume.com")

def call(method, path, body=None, attempts=4):
    data = json.dumps(body).encode() if body is not None else None
    headers = {"x-api-key": os.environ["SUME_API_KEY"]}
    if data is not None:
        headers["content-type"] = "application/json"
        headers["idempotency-key"] = str(uuid.uuid4())  # one key for every attempt
    for attempt in range(attempts):
        req = urllib.request.Request(BASE + path, data=data, method=method, headers=headers)
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                return json.load(resp)
        except urllib.error.HTTPError as exc:
            err = json.load(exc).get("error", {})
            draining = exc.code == 503 and err.get("code") == "deploy_draining"
            if not draining or attempt == attempts - 1:
                raise
            wait = err.get("retry_after_seconds") or int(exc.headers.get("retry-after", 5))
            time.sleep(wait)

What to log

Log error.request_id on every failed attempt. If four attempts in a row return deploy_draining, which is about 15 seconds of waiting with the default delay, stop and surface the error. A drain is normally over well before that. For a long-running generation, the job state lives on the server, so poll its status_url instead of resubmitting.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume