fal low priority and start_timeout vs Sume queue-first admission

fal queues accept a low priority and a start_timeout that 504s. Sume has neither: it queues by plan and rejects with queue_full. What to do when work must wait.

5 min readSume
All posts

Can I set a priority or a start deadline on a Sume job?

No. Sume's generation docs do not list a priority field or a start deadline on submit. A job waits in queued until your workspace has a free processing slot, and your plan decides how many jobs can wait. fal's queue API, by contrast, takes a priority of normal or low and a start_timeout.

What do fal's two parameters do?

From fal's queue documentation: priority accepts "normal" (default) or "low", and low-priority requests queue behind normal requests on the same endpoint. start_timeout is a server-side absolute deadline in seconds; if it passes before processing begins, fal returns 504 Gateway Timeout.

fal also describes IN_QUEUE as received and stored while waiting for a runner, and says requests in that state are removed immediately on cancel. Together these give you a way to say "run this when there is room, and give up if it never starts".

How does Sume treat waiting work instead?

Sume uses queue-first admission. A valid submit with reservable balance is accepted as queued even if the workspace is at its concurrency limit, and workers move jobs to processing under a per-workspace guard. Concurrency is plan-only; top-ups do not raise it. See Generation admission.

The queue is bounded. Default queue capacity is max(3, concurrency_limit x 5), and when accepted capacity is full the submit fails with 429 queue_full. Request-rate abuse limits fail separately with 429 rate_limited. Cancellation works only before generation starts; after that the API returns 409 job_generation_already_started.

Queue controls, read 2026-10-02
PlanProcessing concurrencyQueue capacity (default)Accepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

How do I reproduce a start deadline on Sume?

You do it in your client. Submit with async and an Idempotency-Key, poll status_url, and if the job is still queued past your deadline, call POST /v1/jobs/:id/cancel. That is safe only while the job has not started; if it returns 409 job_generation_already_started, the job runs to completion and bills.

Do not use wait_timeout_seconds for this. It bounds how long the submit HTTP call blocks (0 to 30 seconds) and never changes how long a job may run or wait. A client-side timeout does not cancel anything.

Which model of waiting fits your workload?

fal's low priority suits background batches that should yield to interactive traffic on one endpoint. Sume's approach suits steady pipelines where you want predictable admission: read generation_limits on each submit response, size your wave from queue_capacity_remaining, and treat queued as normal.

  • Need strict ordering between interactive and batch work? Sume has no priority lane documented, so run two workspaces or pace the batch yourself.
  • Need a hard deadline? Poll and cancel before start.
  • Retrying a refused submit? Reuse the same Idempotency-Key so a retry returns the original job.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume