Does a queue_full 429 charge me? Sume's reservation rules

A Sume 429 queue_full means the workspace has no accepted-job capacity left. The failed admission releases its reservation; retry with the same Idempotency-Key.

5 min readSume
All posts

A 429 queue_full from Sume means your workspace has used all its accepted generation capacity, so no new paid job can be admitted. The docs say Sume releases or refunds the reservation for the failed admission where applicable, so the rejected submit should not leave a charge behind. Wait for capacity, then retry with the same Idempotency-Key.

This page covers what the error does to your balance and your queue, using Generation admission and Errors and rate limits.

What does queue_full actually mean?

Sume admits paid generation queue-first. A valid submit creates a durable job when balance can be reserved and the workspace still has accepted-job capacity. Concurrency being full is not an error by itself: extra jobs wait as queued. The error appears only when processing slots and queue are both full.

Accepted job capacity is the processing concurrency limit plus the queue limit, and the queue default is max(3, concurrency_limit x 5). The effective numbers are in generation_limits on submit responses, and the docs tell you to prefer that field over the static plan table.

Default capacity by plan (read 2026-10-03)
PlanProcessingQueueAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Is anything charged when the submit is rejected?

At submit Sume reserves the estimated amount when it accepts a request. Successful completion captures the reservation; failed jobs and failed queue admission release or refund it where applicable. A queue_full response can carry a generation_limits snapshot and metadata about the failed admission attempt.

The docs hedge with where applicable, so verify on your own account instead of assuming: read GET /v1/balance before a burst and after, and compare. The related 402 insufficient_credits is the opposite case: Sume could not reserve the estimate at all, so no provider work starts.

How is it different from rate_limited?

Both are 429, so branch on error.code. rate_limited is request volume against an abuse-protection limit and carries back-off headers. queue_full is about workspace capacity and clears when an existing job finishes or is canceled.

Do not retry paid submits without an Idempotency-Key in either case. A retry that reuses the key returns the original job instead of billing a second one, which matters if your first attempt was accepted but the response was lost.

What should the client do?

The docs list four steps: stop adding work for that workspace, poll existing jobs until one reaches a terminal state, cancel queued jobs you no longer need, then retry with the same key after capacity opens, honoring retry-after when present. Cancel works only before generation starts; after that it is 409 job_generation_already_started.

The sketch below does the retry part. It waits on retry-after when the header is present and otherwise grows its own delay, and it resends the same key every time.

import asyncio
import os
import httpx

async def submit(body: dict, key: str) -> dict:
    headers = {
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": key,
    }
    delay = 5.0
    async with httpx.AsyncClient() as http:
        for _ in range(8):
            res = await http.post(
                "https://api.sume.com/v1/image-1.0/generate",
                json=body, headers=headers,
            )
            if res.status_code != 429:
                return res.json()
            code = res.json().get("error", {}).get("code")
            if code != "queue_full":
                raise RuntimeError(code)
            wait = float(res.headers.get("retry-after", delay))
            await asyncio.sleep(wait)
            delay = min(delay * 2, 60)
    raise RuntimeError("queue_full after 8 attempts")

print(asyncio.run(submit({"prompt": "matte bottle on marble", "mode": "async"}, "hero-001")))

How do I avoid hitting it?

Size waves from the live snapshot. The docs give a budget for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. wave_size_hint is only a submission-wave hint that includes queue slots; it is not a concurrency limit, and its minimum value is 1 even when the queue is full.

Sume does not expose a per-job queue position or an ETA, and queue expiration is not a public option, so your own pacing is the only lever.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume