fal never drops queued requests; Sume can answer 429 queue_full

fal's queue docs say queued requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full when full. Plan for the gap.

5 min readSume
All posts

fal's queue documentation, read 2026-10-05, says requests in the queue are never dropped: if no runner is free, the request waits while fal scales up. Sume's admission model is different on purpose. Valid jobs are accepted as queued while accepted capacity remains, but when the workspace has no remaining queue capacity, a new paid submit fails with 429 queue_full. A client written for an unbounded queue will treat that response as a bug. It is the designed backpressure signal, and the fix is waiting and retrying with the same idempotency key.

Four controls, four different errors

Sume separates generation concurrency (jobs in processing), queue capacity (accepted jobs not yet started), submit rate limits and wallet balance. Full concurrency alone is not an error, since jobs queue. Only a full queue returns queue_full. A request-volume breach returns 429 rate_limited, and a short wallet returns 402 insufficient_credits before any provider work starts.

What happens when a limit is full (read 2026-10-05)
Limitfal queue pageSume
Queue sizeNo queue size limit, requests never droppedAccepted capacity per plan; 429 queue_full beyond it
Running capacityRunners scale up automaticallyPlan-only concurrency; extra jobs wait as queued
Failed attemptRe-queued, retried up to 10 times on 503, 504 or a connection errorYour client retries; the same idempotency key prevents a double charge
Request volumeNot covered on the fetched page429 rate_limited with retry-after

A client that handles both

Treat 429 as retry-later in every case, but read the code. For queue_full, wait until a job finishes or cancel queued work, then retry the same submit with the same Idempotency-Key. For rate_limited, wait the retry-after seconds. Never retry an unsafe submit without a key, because a response lost on the wire hides a job that was in fact accepted.

import time


def wait_seconds(status, code, retry_after, attempt):
    if status != 429:
        return None
    if code == "rate_limited" and retry_after:
        return retry_after
    return min(60, 2 ** attempt)


for n, args in enumerate([(429, "rate_limited", 3), (429, "queue_full", None)]):
    print(wait_seconds(*args, n))

Why the difference helps

A bounded queue makes cost and latency predictable: you cannot accidentally stack thousands of paid jobs behind a small plan. The price is that your submitter must be a loop, not a fire-and-forget burst.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume