queue_full 429 on a Sume submit: the reservation is released
A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.

A 429 queue_full on a Sume generation submit does not leave you paying for a job that was never accepted. The docs say that, where applicable, Sume releases or refunds the reservation of a failed admission. In the usage ledger that shows up as a refunded row and a non-zero refunded_usd_micros, which the summary labels as not spend.
The right reaction is to stop adding work, wait for capacity, and retry the same request with the same Idempotency-Key. A new key would be a new paid intent.
What queue_full means
queue_full says the workspace used all of its accepted generation capacity. Concurrency alone is not an error: while queue capacity remains, Sume accepts valid jobs as queued. The submit fails only when the processing slots and the queue are both full. The 429 rate_limited code is a different problem, about request volume, and its fix is a backoff rather than waiting for jobs to finish.
The error body carries a code, a message and a request_id. Its details can include a generation_limits snapshot and job metadata for the failed admission attempt. Treat the counts in that snapshot as a point-in-time reading, since workers claim jobs and other clients submit work right after the response.
Two gates, two different money outcomes
The table compares the early rejections for paid submits, from Generation admission.
| Status and code | Stage | Money |
|---|---|---|
402 insufficient_credits | Sume cannot reserve the estimate from the balance | Nothing is reserved; fails before provider work |
429 queue_full | Workspace has no accepted generation capacity left | Reservation of the failed admission is released or refunded where applicable |
429 rate_limited | Request volume over an abuse-protection limit | Not a generation admission failure; use retry-after |
503 provider_capacity_exceeded | Sume cannot dispatch work safely | Retry later with the same key unless told not to |
Retry the same key
The sample submits an Image 1.0 job in async mode. On queue_full it waits for retry-after when the header exists and otherwise ten seconds, which is my own choice because the docs give no default, and sends the identical request with the identical key. Any other 429 stops the loop, because that is the rate-limit case and needs a different handling. It reads the key from the environment.
const base = process.env.SUME_API_BASE_URL ?? "https://api.sume.com/v1";
const key = process.env.SUME_API_KEY;
if (!key) throw new Error("SUME_API_KEY is not set");
const idem = process.argv[2] ?? "hero-2026-10-07-001";
for (let attempt = 1; attempt <= 5; attempt++) {
const res = await fetch(`${base}/image-1.0/generate`, {
method: "POST",
headers: { "x-api-key": key, "content-type": "application/json", "idempotency-key": idem },
body: JSON.stringify({ prompt: "Matte black bottle on marble", mode: "async" }),
});
const body = await res.json();
if (res.status !== 429) {
console.log(res.status, body.request_id ?? body.error?.code);
break;
}
if (body.error?.code !== "queue_full") throw new Error(`rate_limited: back off, then retry`);
const wait = Number(res.headers.get("retry-after") ?? 10);
console.log(`queue_full, attempt ${attempt}, waiting ${wait}s with the same key`);
await new Promise((r) => setTimeout(r, wait * 1000));
}Verify the refund instead of assuming it
If you need proof for finance, query GET /v1/usage with a job_id when the error details include job metadata, or with the thread_id or run_id you hold, and read the summary. A refunded hold keeps its amount in billable_amount_usd_micros, so look at refunded_usd_micros and at final, not at that field. The usage summary post shows the fields.
While the queue is full
- Do not add more work for that workspace.
- Poll current jobs until at least one reaches a terminal state.
- Cancel queued jobs you no longer need. Cancel works only before generation starts.
- Keep
retry-afteras the pacing source when it is present. - Do not resubmit only because a local worker timed out. A timeout on your side is not an admission failure, and the job may already exist.
- Log
request_idfrom each error envelope, so a support ticket can point at the exact attempt. - Pace the next batch from
generation_limitsinstead of firing everything at once.
Sources
Related posts
More in Developers
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
- Redact faces and license plates: Pillow first, AI edit only to replace
For redaction use Pillow boxes you control; use an AI mask edit on openai/gpt-image-2.5 only to replace a plate or face, from $0.0094 per image on Sume.
- Redeliver a missed video webhook after a bad deploy: one Sume call
Receiver down when the video finished? POST /v1/jobs/{job_id}/webhook/redeliver re-sends job.completed with a fresh signature. Scope, statuses, pitfalls.
- Restyle avatar clip captions with source_caption_id, no re-transcribe
To try a second caption look on an avatar video, send source_caption_id instead of the video URL. Sume reuses the word timings. Cost, errors and a worked flow.
Written by Sume