Circuit breaker pattern in Python for AI API calls

A circuit breaker stops calling a failing API: after repeated failures it opens and fails fast, then lets a trial call through. A Python version.

5 min readSume
All posts

The circuit breaker pattern stops your code from calling a service that keeps failing. In Python it is a small wrapper around the call: it counts recent failures, and past a threshold it opens and raises at once instead of calling; after a cooldown it lets a trial call through, and closes again if that call succeeds.

The pattern's states and rules below come from Microsoft's Circuit Breaker pattern page, read on 2026-09-28. The error codes that decide what counts as a failure come from Sume's Errors and rate limits docs, as an example of a paid AI generation API.

What are the three circuit breaker states?

Microsoft describes the breaker as a proxy that works like a state machine with three states:

  • Closed: calls go through. The proxy counts recent failures, and when they pass a threshold within a time period it switches to Open and starts a timer.
  • Open: calls fail immediately with an exception. The API is not called.
  • Half-Open: when the timer expires, a limited number of calls pass through. If they succeed, the breaker closes and resets its failure counter. If any fails, it reopens and restarts the timer.
  • The Closed-state counter is time based, so occasional failures spread over a long period don't open the breaker.

How do I write a circuit breaker in Python?

This version keeps a sliding window of failure times. A network error or a 5xx response is a failure; a 4xx response is not, because it is about your request, not the service. After the cooldown, the next call is the half-open trial.

import time
import requests


class CircuitBreaker:
    def __init__(self, threshold=5, window=60.0, cooldown=30.0):
        self.threshold, self.window, self.cooldown = threshold, window, cooldown
        self.failures: list[float] = []
        self.opened_at: float | None = None  # None = closed

    def call(self, send):
        now = time.monotonic()
        if self.opened_at is not None and now - self.opened_at < self.cooldown:
            raise RuntimeError("circuit open: not calling the API")
        resp = None
        try:
            resp = send()  # after the cooldown, this is the half-open trial
        finally:
            if resp is None or resp.status_code >= 500:  # network error or 5xx
                self.failures = [t for t in self.failures if now - t < self.window] + [now]
                if self.opened_at is not None or len(self.failures) >= self.threshold:
                    self.opened_at = now  # open, or reopen after a failed trial
            elif self.opened_at is not None:
                self.failures, self.opened_at = [], None  # trial passed: close
        return resp

How do I use it for AI API calls?

Share one breaker per dependency across your workers, and send each submit with an Idempotency-Key: on Sume's POST /v1/videos, a replay with the same key returns the original job (Video generation). That includes a refused one: in current Sume code, a same-key retry after provider_capacity_exceeded or queue_full replays that error, so once the breaker closes, resubmit those under a new key. Microsoft notes that many instances may use the same breaker, so it must not block concurrent calls; in threaded code, guard the state with a lock.

import os

breaker = CircuitBreaker()


def submit(prompt: str, key: str) -> requests.Response:
    return breaker.call(lambda: requests.post(
        "https://api.sume.com/v1/videos",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
                 "Idempotency-Key": key},
        json={"model": "seedance-2", "prompt": prompt},
        timeout=30,
    ))

Which API errors should trip the breaker?

Only the ones that say the service is unwell. A breaker that opens on your own bad requests, or on your account being busy, stops good work for the wrong reason. Microsoft's page separates the two signals: 429 when a service throttles the client, 503 when the service isn't available.

From Sume's Errors and rate limits, read 2026-09-28.
ResponseWhat Sume says it meansCount as a failure?
503 provider_capacity_exceededSume's provider dispatch queue is fullYes
503 provider_not_configuredProvider execution is unavailable; do not retry aggressivelyYes
429 rate_limitedToo many requests in the current window; use retry-after when presentNo: back off
429 queue_fullWorkspace generation concurrency plus queue capacity is fullNo: submit less
400, 401, 402Invalid request, bad API key, or balance too lowNo: fix the request

How do retry and circuit breaker patterns work together?

They do different jobs. A retry expects the call to succeed eventually; a breaker stops calls that are likely to fail. Microsoft suggests combining them: retry through the breaker, and stop retrying when the breaker says the fault is not transient.

In the code above, that means catching the RuntimeError in your retry loop and ending the loop instead of sleeping and trying again. Python requests retry with backoff covers the retry half.

  • queue_full is not an outage: Sume's docs say it means another paid job can't be accepted until a queued or processing job finishes or is canceled. Lower your concurrency instead.
  • Failed jobs carry an error category too. runtime_unavailable says "retry later; do not retry aggressively", which a second breaker around resubmits can honor (Errors and rate limits).
  • In a queue-driven pipeline, Microsoft notes that failed messages often go to a dead-letter queue instead; see What is a dead letter queue?.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume