fal never drops queued requests; Sume can answer 429 queue_full
fal's queue docs say queued requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full when full. Plan for the gap.

fal's queue documentation, read 2026-10-05, says requests in the queue are never dropped: if no runner is free, the request waits while fal scales up. Sume's admission model is different on purpose. Valid jobs are accepted as queued while accepted capacity remains, but when the workspace has no remaining queue capacity, a new paid submit fails with 429 queue_full. A client written for an unbounded queue will treat that response as a bug. It is the designed backpressure signal, and the fix is waiting and retrying with the same idempotency key.
Four controls, four different errors
Sume separates generation concurrency (jobs in processing), queue capacity (accepted jobs not yet started), submit rate limits and wallet balance. Full concurrency alone is not an error, since jobs queue. Only a full queue returns queue_full. A request-volume breach returns 429 rate_limited, and a short wallet returns 402 insufficient_credits before any provider work starts.
| Limit | fal queue page | Sume |
|---|---|---|
| Queue size | No queue size limit, requests never dropped | Accepted capacity per plan; 429 queue_full beyond it |
| Running capacity | Runners scale up automatically | Plan-only concurrency; extra jobs wait as queued |
| Failed attempt | Re-queued, retried up to 10 times on 503, 504 or a connection error | Your client retries; the same idempotency key prevents a double charge |
| Request volume | Not covered on the fetched page | 429 rate_limited with retry-after |
A client that handles both
Treat 429 as retry-later in every case, but read the code. For queue_full, wait until a job finishes or cancel queued work, then retry the same submit with the same Idempotency-Key. For rate_limited, wait the retry-after seconds. Never retry an unsafe submit without a key, because a response lost on the wire hides a job that was in fact accepted.
import time
def wait_seconds(status, code, retry_after, attempt):
if status != 429:
return None
if code == "rate_limited" and retry_after:
return retry_after
return min(60, 2 ** attempt)
for n, args in enumerate([(429, "rate_limited", 3), (429, "queue_full", None)]):
print(wait_seconds(*args, n))Why the difference helps
A bounded queue makes cost and latency predictable: you cannot accidentally stack thousands of paid jobs behind a small plan. The price is that your submitter must be a loop, not a fire-and-forget burst.
Sources
Related posts
More in Developers
- Fallback chain for a 30-second AI video: first catalog row that fits
Kling 4.0 is rolling out in stages. Pick the first Sume video id whose catalog row lists 30 seconds, and stop with an error if none does.
- Fan out one Sume clip to three platforms: one webhook, three task keys
Receive one job.completed webhook per Sume job, then enqueue a task per platform. Key each task by job id plus platform so a retry never double-posts.
- Sume video content 409: job_not_completed vs job_failed in Python
A 409 from /v1/videos/{id}/content means two things. job_not_completed is retryable, job_failed is not. Here is a Python handler that tells them apart.
- Fields Omni rejects on Sume: generate_audio false, bitrate_mode
Which request fields Sume's gemini-omni-flash-1.1 refuses or lacks: generate_audio false, bitrate_mode, reference_audio_urls, plus edit-mode rules.
Written by Sume