queue_admission_unavailable: a failed Sume job you can retry

If Sume cannot evaluate generation admission, the job fails with queue_admission_unavailable, retryable true, and is refunded. Read it, then create again.

4 min readSume
All posts

Every paid Sume generation passes an admission check that decides if the job can start now, wait in the queue, or be refused for capacity. Usually you see the answers in the docs: a job runs, or queue_full returns a failed job with retry_after_seconds 30. A third answer exists for the case where the check itself breaks. The job is marked failed with the error code queue_admission_unavailable.

What the failed job says

The error message is "Generation admission could not be evaluated. Retry the request." and the error has retryable set to true. Nothing was sent to a model provider. That is the point of the design: a job whose admission cannot be vouched for is not left sitting in the queue, because a stuck queued row would count against your workspace and hold your funds.

The server also refunds the usage reserved for the job, with the reason queue_admission_unavailable. If the refund itself fails because the same database trouble blocks it, a repair task retries the cleanup once the pool recovers, so the hold is not kept forever.

How to handle it

Treat it like queue_full, with one difference. queue_full says you are over capacity and gives a 30 second hint. This code says Sume could not tell. Wait a few seconds with jitter and create again. Read the job first (GET /v1/jobs/{id}) and confirm status failed and the error code, so you act on the real cause.

On the retry, think about your idempotency key. A replay of a key returns the original response, so if you reuse the key you may be handed the same failed job back. Derive a fresh key per attempt, such as the order id plus an attempt counter, and keep the attempt count small.

import os, time, random, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def retryable_failure(job):
    err = (job.get("error") or {})
    return job.get("status") == "failed" and err.get("retryable") is True

def backoff(attempt):
    time.sleep(min(30, 2 ** attempt) + random.random())

What it is not

It is not a rate limit on your key, so the ratelimit headers will look normal. It is not a content refusal, because those are not retryable. And it is not a reason to lower your spend cap or concurrency. Alert on it as a platform-side signal: a single one is noise, a run of them is worth a support note with the job ids.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume