fal queue has no size limit: what Sume does instead (queue_full)
fal's queue docs say there is no queue size limit and requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full. What that means.

fal's asynchronous inference page says the queue has no size limit and that requests in it "are never dropped". Sume takes the opposite design choice: each workspace has a fixed number of accepted generation jobs (processing plus queued), and a submit beyond that returns 429 queue_full instead of piling up.
Neither choice is wrong. They move the backpressure to different places, and that changes how you write the client. This page reads fal's queue page as published on 2026-10-02 and Sume's own docs, and only states what each says.
What does fal say about its queue?
On the fal queue page, a request submitted to https://queue.fal.run/{model-id} is stored and moves through three statuses: IN_QUEUE (received and stored, waiting for a runner), IN_PROGRESS (routed to a runner) and COMPLETED (result stored). The page also states that fal scales automatically as demand changes.
The submit call also takes priority (normal by default, or low) and start_timeout, a server-side deadline that returns a 504 if processing has not started in time. Those are the client-side tools fal gives you for a long queue. Because the queue itself has no size cap, a flood of submissions is accepted and simply waits.
What does Sume do when you submit too much?
Sume uses queue-first admission. A valid paid job is accepted as queued when the balance can be reserved and the workspace still has accepted-job capacity. Concurrency is a dispatch limit, not a submit limit, so being at your processing cap is not an error by itself.
The cap shows up when the queue fills. Queue capacity defaults to max(3, concurrency_limit x 5), and the docs list these defaults:
- Concurrency is plan-only: prepaid top-ups do not raise it.
- The effective numbers come back in
generation_limitson submit responses, so prefer that field over this table. - When accepted capacity is gone, the submit returns
429 queue_full, the reservation for that attempt is released, and you retry with the sameIdempotency-Key.
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Which design is better for a batch job?
With an unbounded queue you can submit 10,000 items in a loop and walk away; the cost is that you learn about the backlog only by polling status, and a backlog you no longer want is something you must cancel request by request. fal documents start_timeout precisely so a request does not wait forever.
With a bounded queue you find out at submit time. A queue_full response is a clear signal to stop adding work, wait for a job to finish, or cancel queued jobs you no longer need. The downside is that your client must pace itself. Sume's docs give a sizing rule: new in-flight work is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. The wave_size_hint field is only a submission hint, not a concurrency limit.
One more honest limit: Sume currently exposes queue counts and remaining capacity, not a per-job queue position or ETA, and queue expiration is not a public option. If you need a fail-fast deadline on queue time, Sume does not offer a start_timeout equivalent today.
How should the client handle queue_full?
Treat queue_full as pacing feedback, not a failure of the job. The docs say to stop adding generation work for that workspace, poll existing jobs until one is terminal, cancel queued jobs you no longer need, and resubmit with the same idempotency key, honoring retry-after when present. Cancellation works only before generation starts; after that it returns 409 job_generation_already_started.
Keep this separate from 429 rate_limited, which is request volume. Sume budgets reads and writes separately (the Free plan gets 120 writes and 4,800 reads per minute), so a tight status-poll loop cannot rate-limit your own submits.
curl -i -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: batch-42-item-0007" \
-d '{"model":"sume/auto","prompt":"A product clip on a desk, natural light","duration":5}'
# 429 with error.code queue_full: wait, then repeat this exact call
# with the same Idempotency-Key. A replay returns the original job.Which should you pick?
If your workload is bursty and you want the platform to absorb it without a client-side limiter, fal's documented design fits. If you want hard feedback at submit time and a predictable per-plan ceiling on spend and backlog, Sume's bounded queue fits better. Both give you async jobs, status polling and webhooks; the difference is where "too much" is reported.
For the rest of the comparison, see Sume vs fal and fal queue priority and start_timeout vs Sume admission. Full Sume details are in Generation admission.
Sources
Related posts
More in Comparisons
- fal queue_position and logs vs Sume job status, events, usage
fal returns queue_position, runner logs and inference_time. Sume gives status, an events timeline and per-job usage. Which one tells you why a job is slow?
- fal webhook redirect 3xx is never retried: what Sume does instead
fal treats a 3xx from your webhook URL as a permanent failure. Sume does not follow redirects either, but counts a 3xx as a failed attempt, not an end.
- fal webhook ED25519 and JWKS vs Sume's HMAC-SHA256 check
fal signs webhooks with ED25519 keys fetched from a JWKS URL. Sume signs HMAC-SHA256 over timestamp.body with a workspace secret. The two checks, side by side.
- fal webhook retries: 31 attempts, 15 s timeout, vs Sume's 10
fal retries a failed webhook with backoff up to 31 times while the stored result lasts; Sume makes up to 10 attempts with a 10 second timeout. What to build.
Written by Sume