Sume generation_capacity_exhausted: the 503, job reason and flag
Sume reports provider capacity three ways: HTTP 503 provider_capacity_exceeded, a job reason generation_capacity_exhausted, and a sync flag. All three retry.

generation_capacity_exhausted is the public reason Sume gives when the generation provider has no room. It shows up under three names: an HTTP 503 with error.code provider_capacity_exceeded, a failed job whose public_reason is generation_capacity_exhausted, and a capacity flag on a synchronous wait. All three mean wait and retry, not change the input.
What does the HTTP 503 look like?
A submit that hits provider capacity returns 503 provider_capacity_exceeded. The envelope says category: queue, stage: generation_submit, retryable: true, public_reason: generation_capacity_exhausted and next_action: retry_later. retry_after_seconds comes from details.retry_after_seconds when the provider supplied one and is 30 otherwise, rounded up and capped at 300. Like other submit failures it includes the job id in details.
What does the failed job look like?
If the job itself records the capacity failure, the public job error uses category: generation_unavailable, the same stage: generation_submit, retryable: true and the same public reason. Its retry_after_seconds can be null, so do not index into it blindly. Use your own default.
| Where | Name | Category | Retry hint |
|---|---|---|---|
| HTTP response | provider_capacity_exceeded (503) | queue | retry_after_seconds, default 30, max 300 |
| Failed job | generation_capacity_exhausted | generation_unavailable | May be null |
| Sync wait | capacity flag | see the sync explainer | Poll or retry later |
How do I retry safely?
Reuse the same Idempotency-Key for the same intended clip, so a retry cannot create a second paid job, and wait the hinted time. Do not rewrite the prompt; the input was not the problem. This helper returns the wait or None when the error is not a capacity case.
def capacity_wait(error: dict) -> int | None:
is_capacity = (
error.get("code") == "provider_capacity_exceeded"
or error.get("public_reason") == "generation_capacity_exhausted"
)
if not is_capacity or not error.get("retryable", False):
return None
hint = error.get("retry_after_seconds")
return min(int(hint), 300) if isinstance(hint, (int, float)) and hint > 0 else 30
print(capacity_wait({"code": "provider_capacity_exceeded", "retryable": True, "retry_after_seconds": 30}))Sources
Related posts
More in Developers
- Sume job failed artifact_too_large: shrink the output, don't rerun
A Sume job failing with artifact_too_large or artifact_upload_rejected made its file but could not store it. Make the file smaller; rerunning changes nothing.
- pending_usd_micros vs held: what Sume's settle sweeper still owns
Sume /v1/usage splits open holds into held and pending_usd_micros. Read the two fields and settle_state to tell parked rows from spend, and when final flips.
- unpriced_usd_micros: legacy rows still inside Sume's debited total
unpriced_usd_micros in Sume /v1/usage is captured spend from legacy rows with no price-book stamp. It is inside debited, so never subtract or add it twice.
- TanStack Start webhook route for Sume jobs: verify the raw body
Receive a signed Sume job webhook in a TanStack Start server route: read request.text(), verify with the SDK, answer 204, and dedupe on job_id.
Written by Sume