OpenAI Agents Python MCP backoff ceiling vs Sume retry-after
Set the MCP retry backoff ceiling low enough that a Sume 429 with retry-after is honored first, and never retry a paid create without its idempotency key.

The openai-agents-python v0.20.0 notes add a configurable retry backoff ceiling for MCP and derive streamable HTTP retry backoff from the backoffs already taken. Against Sume, the ceiling is a cap on your own waiting, not a replacement for the server's hint: when a 429 carries retry-after, wait at least that long, and treat queue_full as a different case from ordinary rate limiting.
The release page says only that these changes exist; it does not list option names in the part read on 2026-10-01, so check the SDK docs for the exact parameter. Sume behavior is from Errors and credits and Jobs and results.
What does Sume say to do on a 429?
The docs are short: "Back off when you receive 429. Use retry-after when present. Do not retry unsafe submit requests without an Idempotency-Key." Responses can also carry ratelimit-limit, ratelimit-remaining and ratelimit-reset. See rate limit headers for what each one means.
Is queue_full the same as a rate limit?
No. Both return 429, but they have different codes. rate_limited means too many requests in the current window. queue_full means the workspace's generation concurrency plus queue capacity is full, so Sume cannot accept another paid generation job until a queued or processing job finishes or is canceled. Being at concurrency by itself is not an error: valid jobs are accepted as queued while queue capacity remains. A growing exponential backoff does not free a full queue, so branch on the code field rather than the status alone. The longer explanation is in queue_full vs concurrency full.
How should a ceiling interact with the server hint?
| Response | What the docs say | Client behavior |
|---|---|---|
429 rate_limited with retry-after | Use retry-after when present | Wait at least that long, even if your ceiling is lower |
429 queue_full | Needs an existing job to finish or be canceled | Stop submitting new paid jobs; read job status first |
provider_capacity_exceeded | Retry later with the same idempotency key | Backoff is fine; reuse the key |
| Submit retry | Reuse the same Idempotency-Key | The retry returns the original job instead of billing twice |
What about waits on a slow job?
Backoff is for failed requests. A job that is simply still running should be waited on with jobs_wait: on remote MCP timeout_seconds defaults to 50 and is capped at 55, and wait_slice_expired means call jobs_wait again with the same ids, never resubmit the create. On the REST side, follow next_poll_after_seconds when present and back off otherwise. See jobs_wait for long video jobs.
Sources
Related posts
- 429 queue_full vs rate_limited: what the generation API means
- Rate limit headers: what limit, remaining and reset mean
- 429 rate_limited: how to tell a read budget from a write budget
- MCP tool call timeouts on long-running video jobs: use jobs_wait
- Axios retry: retry a POST safely with an idempotency key
More in Developers
- Does an OpenAI API key expire? Expiry dates and key rotation
OpenAI project keys can now carry an expiration date and orgs can cap lifetime. Sume documents no expiry field: rotate by minting a replacement key.
- OpenAI async tool calling for long-running render jobs
OpenAI's async tool calling lets the model keep working while your tool runs. For a slow Sume render, return the job id fast, then wait in slices.
- OpenAI image-encoding fix: rerun workflows with the right key
OpenAI fixed an image-encoding bug and advised retrying affected workflows. On Sume, a replayed idempotency key returns the original; use a new key to rerun.
- mTLS instead of an API key? How Sume authenticates calls
OpenAI's docs list mutual TLS and workload identity federation. Sume authenticates API calls with one API key header and signs webhooks with HMAC.
Written by Sume