Tenacity retries forever with no wait: safe Sume submit

Tenacity's default is to retry forever without waiting. Four settings turn it into a bounded, jittered retry that is safe around a paid Sume submit.

5 min readSume
All posts

Short answer

A bare @retry from Tenacity retries forever and does not wait between attempts, per the Tenacity docs. Around a paid Sume submit that is a tight loop of duplicate requests. Add stop_after_attempt, wait_random_exponential, a narrow retry_if_exception_type and reraise=True, and create the Idempotency-Key outside the decorated function.

The four settings

Each setting below comes from the Tenacity documentation. Together they bound the loop in time, spread the attempts out and keep the final error visible.

The key placement matters as much as the decorator. If the function creates its own UUID, every attempt is a new request to Sume. Pass the key in, so every attempt carries the same value and a replay returns the original instead of a second charge.

Tenacity settings that matter for a paid POST (read 2026-10-03)
SettingWhat it does
Default (no arguments)Retries forever, no waiting
stop_after_attempt(n)Gives up after n attempts
wait_exponential / wait_random_exponentialExponential wait of 2^x times a multiplier, with min and max bounds; the random variant adds jitter
retry_if_exception_typeRetries only on the listed exception types
reraise=TrueRaises the last underlying error, not a RetryError

What to retry on a Sume call

Retry transport errors, 408, 429 and 5xx. The Sume client library uses that same set. A 429 can be rate_limited or queue_full; both are worth retrying with the same key, and the response carries retry-after. Do not retry 400, 401, 402, 403 or 404, because the same request gets the same answer.

A 409 needs a look at the code. idempotency_key_in_use is retryable, since another request with the same key is still running. idempotency_conflict is not: you reused a key with a different body, and no amount of waiting fixes that. The full list is on the errors page.

import os
import uuid

import httpx
from tenacity import (retry, retry_if_exception_type, stop_after_attempt,
                      wait_random_exponential)


class Retryable(Exception):
    pass


@retry(retry=retry_if_exception_type((Retryable, httpx.TransportError)),
       stop=stop_after_attempt(4),
       wait=wait_random_exponential(multiplier=1, max=20),
       reraise=True)
def submit(key: str, body: dict) -> dict:
    r = httpx.post("https://api.sume.com/v1/image-1.0/generate", json=body,
                   headers={"x-api-key": os.environ["SUME_API_KEY"],
                            "Idempotency-Key": key}, timeout=30)
    if r.status_code in (408, 429) or r.status_code >= 500:
        raise Retryable(r.status_code)
    r.raise_for_status()
    return r.json()


if __name__ == "__main__":
    print(submit(str(uuid.uuid4()),
                 {"prompt": "A matte black bottle on marble", "mode": "async"}))

After the submit

A successful async submit returns 202 with a job id and a status URL. Retrying the submit is safe with the same key. Polling is a separate loop and needs its own bounds; see Python httpx submit and poll. If your process dies after a timeout, remember that a client timeout does not cancel the job.

Testing the decorator without a network

Tenacity wraps the function, so you can test the policy by pointing the function at a fake transport. With httpx you can pass a mock transport, and the fixture can return 503 twice and 202 on the third attempt. Assert that the same Idempotency-Key header appears on all three requests. That one assertion catches the bug that matters, a key minted inside the retried function.

Also test the failure ends: return 503 on every attempt and confirm that the call raises the original error after four tries rather than looping. With reraise set to True, the exception you catch is the real one, so your logging and alerts see the HTTP status instead of a wrapper type.

Where it fits next to the SDK

If you use the TypeScript SDK you get this policy built in: 2 retries by default, an exponential backoff capped at 8 seconds, and POST only with a key. In Python you assemble the same policy yourself, and Tenacity is a clean way to do it. Keep the retry count low. Each attempt of a submit that reached the server is the same logical request, so more attempts rarely help and mostly delay the error you need to see.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume