queue_full 429 on a Sume submit: the reservation is released

A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.

5 min readSume
All posts

A 429 queue_full on a Sume generation submit does not leave you paying for a job that was never accepted. The docs say that, where applicable, Sume releases or refunds the reservation of a failed admission. In the usage ledger that shows up as a refunded row and a non-zero refunded_usd_micros, which the summary labels as not spend.

The right reaction is to stop adding work, wait for capacity, and retry the same request with the same Idempotency-Key. A new key would be a new paid intent.

What queue_full means

queue_full says the workspace used all of its accepted generation capacity. Concurrency alone is not an error: while queue capacity remains, Sume accepts valid jobs as queued. The submit fails only when the processing slots and the queue are both full. The 429 rate_limited code is a different problem, about request volume, and its fix is a backoff rather than waiting for jobs to finish.

The error body carries a code, a message and a request_id. Its details can include a generation_limits snapshot and job metadata for the failed admission attempt. Treat the counts in that snapshot as a point-in-time reading, since workers claim jobs and other clients submit work right after the response.

Two gates, two different money outcomes

The table compares the early rejections for paid submits, from Generation admission.

Status and codeStageMoney
402 insufficient_creditsSume cannot reserve the estimate from the balanceNothing is reserved; fails before provider work
429 queue_fullWorkspace has no accepted generation capacity leftReservation of the failed admission is released or refunded where applicable
429 rate_limitedRequest volume over an abuse-protection limitNot a generation admission failure; use retry-after
503 provider_capacity_exceededSume cannot dispatch work safelyRetry later with the same key unless told not to

Retry the same key

The sample submits an Image 1.0 job in async mode. On queue_full it waits for retry-after when the header exists and otherwise ten seconds, which is my own choice because the docs give no default, and sends the identical request with the identical key. Any other 429 stops the loop, because that is the rate-limit case and needs a different handling. It reads the key from the environment.

const base = process.env.SUME_API_BASE_URL ?? "https://api.sume.com/v1";
const key = process.env.SUME_API_KEY;
if (!key) throw new Error("SUME_API_KEY is not set");
const idem = process.argv[2] ?? "hero-2026-10-07-001";

for (let attempt = 1; attempt <= 5; attempt++) {
  const res = await fetch(`${base}/image-1.0/generate`, {
    method: "POST",
    headers: { "x-api-key": key, "content-type": "application/json", "idempotency-key": idem },
    body: JSON.stringify({ prompt: "Matte black bottle on marble", mode: "async" }),
  });
  const body = await res.json();
  if (res.status !== 429) {
    console.log(res.status, body.request_id ?? body.error?.code);
    break;
  }
  if (body.error?.code !== "queue_full") throw new Error(`rate_limited: back off, then retry`);
  const wait = Number(res.headers.get("retry-after") ?? 10);
  console.log(`queue_full, attempt ${attempt}, waiting ${wait}s with the same key`);
  await new Promise((r) => setTimeout(r, wait * 1000));
}

Verify the refund instead of assuming it

If you need proof for finance, query GET /v1/usage with a job_id when the error details include job metadata, or with the thread_id or run_id you hold, and read the summary. A refunded hold keeps its amount in billable_amount_usd_micros, so look at refunded_usd_micros and at final, not at that field. The usage summary post shows the fields.

While the queue is full

  • Do not add more work for that workspace.
  • Poll current jobs until at least one reaches a terminal state.
  • Cancel queued jobs you no longer need. Cancel works only before generation starts.
  • Keep retry-after as the pacing source when it is present.
  • Do not resubmit only because a local worker timed out. A timeout on your side is not an admission failure, and the job may already exist.
  • Log request_id from each error envelope, so a support ticket can point at the exact attempt.
  • Pace the next batch from generation_limits instead of firing everything at once.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume