got maxRetryAfter and 429: rate_limited vs queue_full on Sume

got retries 429 and honors Retry-After up to maxRetryAfter. Sume sends 429 for rate_limited and for queue_full, which need different waits. Here is the split.

5 min readSume
All posts

The answer

got treats 429 as retryable by default and, per its documentation, applies an upper limit to the Retry-After header called maxRetryAfter, which defaults to the request timeout. If the header asks for longer than that limit, the request is aborted rather than retried.

Sume returns two different 429 codes. rate_limited means too many requests in the current window, and retry-after can help. queue_full means the workspace cannot accept another paid generation job until an existing queued or processing job finishes or is canceled, so waiting a few seconds does not clear it.

How the two 429s differ

Sume's errors docs say to back off on 429 and use retry-after when present, and separately that queue_full is different from ordinary rate limiting. A generic retry policy cannot tell them apart by status code alone, so branch on the error code in the body.

Sume 429 codes and the right client response (read 2026-10-03)
CodeMeaningClient response
rate_limitedToo many requests in the current windowWait for retry-after, then retry with the same Idempotency-Key
queue_fullConcurrency plus queue capacity is full for the workspaceDo not loop; wait for a job to finish or cancel one

Wire it with got hooks

got can throw RetryError from a hook to force a retry, and its calculateDelay option can return 0 to abort. The documentation says returning 0 aborts the retry, which is the lever for queue_full: compute the delay as usual, but return 0 when the body says the queue is full. The snippet below does the check inside calculateDelay, reading the response body carried by the error.

import got from "got";

const key = "mug-hero-2026-10-03-r1";
const res = await got
  .post("https://api.sume.com/v1/images", {
    headers: {
      authorization: `Bearer ${process.env.SUME_API_KEY}`,
      "idempotency-key": key,
    },
    json: { model: "sume/auto", prompt: "A mug", mode: "async" },
    retry: {
      limit: 3,
      methods: ["POST"],
      maxRetryAfter: 30000,
      calculateDelay: ({ computedValue, error }) => {
        const body = String(error.response?.body ?? "");
        return body.includes("queue_full") ? 0 : computedValue;
      },
    },
  })
  .json();
console.log(res.data.request_id);

Choosing maxRetryAfter

Because the default maxRetryAfter equals the request timeout, a server hint longer than your timeout aborts the call. That is a reasonable default for an interactive request and a poor one for a batch worker, where waiting a minute is fine. Set it per worker, and keep the Idempotency-Key constant across the waits so a late success is not billed twice.

For queue_full, the documented way out is to read the workspace's running jobs and let one finish or cancel it; see Sume's generation admission page for the queue behaviour and errors docs for the code table.

What to log

Log the Sume request id from the error body, not the key. The errors docs say error bodies include a request id that is safe to share with Sume support, while API keys, signed URLs and raw media URLs should never be included in a report.

Also log which of the two 429 codes you saw. If most of your 429s are queue_full, the fix is concurrency on your side: submit fewer jobs at once, or wait on a batch before the next wave. If most are rate_limited, spacing the submits is enough. Mixing the two in one metric hides which lever to pull.

got's defaults put 429 in statusCodes and cap the header-driven wait with maxRetryAfter, so neither default is wrong; they are just tuned for a request-response call, not for a paid, queued job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume