503 provider_capacity_exceeded: same idempotency key or new?
A Sume 503 provider_capacity_exceeded means the provider queue is full. Docs say retry later with the same key; code replays the refusal, so use a new key.

503 provider_capacity_exceeded from Sume means the provider dispatch queue is full. It is not a 429: your request rate and your workspace queue are not the cause, so slowing your polling does not help. The docs say to retry later with the same idempotency key, but in current code a same-key retry of a submit that already failed this way answers with the stored refusal again. Wait, then send the request with a new Idempotency-Key.
This follows Errors and rate limits and Generation admission, read 2026-09-29; the same-key replay is current API behavior, not something those pages state.
Which 503s can I retry?
Only some. Sume lists several backpressure and runtime codes, and the client behavior differs by code.
| Code | Meaning | What the docs tell the client |
|---|---|---|
provider_capacity_exceeded | Sume's provider dispatch queue is full | Retry later with the same idempotency key (see the next section) |
provider_not_configured | Provider execution is unavailable in this runtime | Do not retry aggressively; check catalog and runtime status |
job_ledger_not_configured | Job persistence is unavailable | Treat as service unavailable |
Should the retry reuse the same idempotency key?
It depends on what you received. If the request timed out or the connection dropped and you never saw a response, resend the identical request with the same Idempotency-Key: with the same key and body the retry returns the original job instead of billing a second one, and the docs say not to retry unsafe submit requests without a key.
If you did receive 503 provider_capacity_exceeded, the submit already failed. In current code a same-key retry replays that failed job's provider_capacity_exceeded instead of trying again, so waiting does not help while the key stays the same. Wait for capacity, then send a new request with a new key. Reserved money is released where applicable when queue admission fails, so the refused attempt does not keep a hold. 429 queue_full behaves the same way on a same-key retry.
How should the client wait?
Use retry-after when the response has it, and otherwise back off. Keep the attempt count small, and never turn the loop into a hammer against a queue that is already full. This bash loop backs off exponentially and gives each attempt its own key, because each 503 was a completed refusal:
delay=2
for attempt in 1 2 3 4 5; do
code=$(curl -sS -o resp.json -w "%{http_code}" \
-X POST https://api.sume.com/v1/image-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hero-shot-2026-09-29-001-try$attempt" \
-d '{"prompt":"Product hero shot of a matte black bottle on marble","mode":"async"}')
[ "$code" != "503" ] && break
sleep "$delay"; delay=$((delay * 2))
done
cat resp.jsonIs this the same as a failed job?
No. A 503 at submit is a refusal before provider work is accepted. A job that was accepted and later fails exposes public error metadata instead, such as a category and a next action. Two of those categories read like capacity: generation_unavailable says retry later, and runtime_unavailable says retry later without retrying aggressively. Keep the two paths separate in your client: one loop for submit refusals, one for terminal job errors.
Compare it with 429 queue_full, which means the workspace has no room for another accepted paid job until an existing one finishes or is canceled. That one is about your workspace; the 503 is about Sume's provider dispatch.
What if it keeps failing?
Read the code in the body before you retry again. If it is provider_not_configured, stop; that is a runtime state to check, not a wait. Otherwise quote the request_id from the error body when you contact Sume support. It is safe to share, unlike API keys, signed URLs or raw media URLs.
On Format runs the same class of failure appears as 503 studio_agent_upstream_unavailable, a Sume-side outage that the Format API errors page answers with a retry using the same Idempotency-Key.
Sources
Related posts
More in Developers
- Which field do I branch on in a Sume API error: code or next_action?
Branch on the HTTP status, then on error.code. Use retryable, retry_after_seconds and next_action to decide on a resend. Never match on message.
- Sume webhook retry schedule: 10 attempts, then what?
Sume tries a webhook up to 10 times. Job webhooks use a fixed 30 s gap; run webhooks back off with jitter up to an hour. What happens next, and how to replay.
- Suno alternative with an API: what Sume's Music Router does
Looking for a music generator you can call from code? What Sume's Music Router takes in, returns and does not do, so you can decide if it fits.
- Talking avatar in JS: make one from Node, play it in React
A talking avatar in JavaScript: create it and send it a script from Node with the Sume SDK, wait for the job, then play the returned MP4 in React.
Written by Sume