Tenacity retries forever with no wait: safe Sume submit
Tenacity's default is to retry forever without waiting. Four settings turn it into a bounded, jittered retry that is safe around a paid Sume submit.

Short answer
A bare @retry from Tenacity retries forever and does not wait between attempts, per the Tenacity docs. Around a paid Sume submit that is a tight loop of duplicate requests. Add stop_after_attempt, wait_random_exponential, a narrow retry_if_exception_type and reraise=True, and create the Idempotency-Key outside the decorated function.
The four settings
Each setting below comes from the Tenacity documentation. Together they bound the loop in time, spread the attempts out and keep the final error visible.
The key placement matters as much as the decorator. If the function creates its own UUID, every attempt is a new request to Sume. Pass the key in, so every attempt carries the same value and a replay returns the original instead of a second charge.
| Setting | What it does |
|---|---|
| Default (no arguments) | Retries forever, no waiting |
| stop_after_attempt(n) | Gives up after n attempts |
| wait_exponential / wait_random_exponential | Exponential wait of 2^x times a multiplier, with min and max bounds; the random variant adds jitter |
| retry_if_exception_type | Retries only on the listed exception types |
| reraise=True | Raises the last underlying error, not a RetryError |
What to retry on a Sume call
Retry transport errors, 408, 429 and 5xx. The Sume client library uses that same set. A 429 can be rate_limited or queue_full; both are worth retrying with the same key, and the response carries retry-after. Do not retry 400, 401, 402, 403 or 404, because the same request gets the same answer.
A 409 needs a look at the code. idempotency_key_in_use is retryable, since another request with the same key is still running. idempotency_conflict is not: you reused a key with a different body, and no amount of waiting fixes that. The full list is on the errors page.
import os
import uuid
import httpx
from tenacity import (retry, retry_if_exception_type, stop_after_attempt,
wait_random_exponential)
class Retryable(Exception):
pass
@retry(retry=retry_if_exception_type((Retryable, httpx.TransportError)),
stop=stop_after_attempt(4),
wait=wait_random_exponential(multiplier=1, max=20),
reraise=True)
def submit(key: str, body: dict) -> dict:
r = httpx.post("https://api.sume.com/v1/image-1.0/generate", json=body,
headers={"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": key}, timeout=30)
if r.status_code in (408, 429) or r.status_code >= 500:
raise Retryable(r.status_code)
r.raise_for_status()
return r.json()
if __name__ == "__main__":
print(submit(str(uuid.uuid4()),
{"prompt": "A matte black bottle on marble", "mode": "async"}))
After the submit
A successful async submit returns 202 with a job id and a status URL. Retrying the submit is safe with the same key. Polling is a separate loop and needs its own bounds; see Python httpx submit and poll. If your process dies after a timeout, remember that a client timeout does not cancel the job.
Testing the decorator without a network
Tenacity wraps the function, so you can test the policy by pointing the function at a fake transport. With httpx you can pass a mock transport, and the fixture can return 503 twice and 202 on the third attempt. Assert that the same Idempotency-Key header appears on all three requests. That one assertion catches the bug that matters, a key minted inside the retried function.
Also test the failure ends: return 503 on every attempt and confirm that the call raises the original error after four tries rather than looping. With reraise set to True, the exception you catch is the real one, so your logging and alerts see the HTTP status instead of a wrapper type.
Where it fits next to the SDK
If you use the TypeScript SDK you get this policy built in: 2 retries by default, an exponential backoff capped at 8 seconds, and POST only with a key. In Python you assemble the same policy yourself, and Tenacity is a clean way to do it. Keep the retry count low. Each attempt of a submit that reached the server is the same logical request, so more attempts rarely help and mostly delay the error you need to see.
Sources
Related posts
More in Developers
- Test a language on Sume STT with a 30-second sample first
Vendors split languages into trained and verified. On Sume STT, language_code is a hint, so cut a 30-second sample and compare auto-detect with your hint.
- Test your Sume webhook receiver locally: openssl and curl
Sign a fake job.completed body with openssl, post it with curl and confirm your receiver answers 200 for valid, and 401 for stale and forged Sume deliveries.
- Test two caption looks on one Reel with source_caption_id
Restyling captions with source_caption_id reuses the first caption job's video and word timings, so no second transcription runs. Billing stays one render each.
- How to test a webhook URL before a Sume Format run uses it
POST /v1/webhooks/test-deliveries sends a signed webhook.test event to your URL. See the scope, the response fields, and the secret check, with no paid run.
Written by Sume