fal queue_position and logs vs Sume job status, events, usage

fal returns queue_position, runner logs and inference_time. Sume gives status, an events timeline and per-job usage. Which one tells you why a job is slow?

5 min readSume
All posts

What does fal tell you while a request runs?

Three things beyond the state name. fal's queue documentation says status() takes an optional with_logs=True (Python) or logs: true (JavaScript) to return runner output, includes queue_position while a request is queued, and includes metrics.inference_time once complete. Sume's job status gives you the state, a timeline of events, and the cost of the job, but not runner logs.

What does Sume give you for the same questions?

For "where is my job", GET /v1/jobs/:id/status returns the lightweight state with terminal, result_ready and next_poll_after_seconds, plus a queue-shaped status that maps one-to-one onto sume_status. The jobs docs do not describe a per-job queue position on that payload.

For "what happened", GET /v1/jobs/:id/events returns a public timeline: job.created, job.queued, job.started, generation.submitted, job.completed, job.failed, job.canceled and webhook.delivery. Public events do not expose raw provider task ids or URLs. For "how much did it cost", GET /v1/usage?job_id=... sums the ledger rows for that job. See Usage.

Side by side

Observability fields, from each vendor's own docs.

Job observability, read 2026-10-02
QuestionfalSume
Is it waiting?IN_QUEUE, with queue_positionqueued; plan queue counts in generation_limits at submit
What is it doing?Runner logs via with_logs / logs: trueEvent timeline; no runner logs
How long did inference take?metrics.inference_time when completeCompare the job.started and job.completed events
What did it cost?Not on the queue page read/v1/usage?job_id= with debited_usd
Pushed progressStatus stream over SSE or pollingTerminal webhooks only; poll or read events

How do I find out why a Sume job is slow?

Start with the queue, because most slow Sume jobs are waiting rather than running. The submit response carries generation_limits with active_generation_jobs, queued_generation_jobs and queue_capacity_remaining. If your workspace is at its processing limit, new jobs sit in queued by design. Then read the events to see when job.started landed.

# job id comes from the submit response
curl https://api.sume.com/v1/jobs/job_123/events \
  -H "Authorization: Bearer $SUME_API_KEY"

curl "https://api.sume.com/v1/usage?job_id=job_123" \
  -H "Authorization: Bearer $SUME_API_KEY"

What do you lose by moving from fal to Sume?

If you tail runner logs to debug a model, you lose that: Sume does not publish them. If you poll queue_position to drive a progress bar, replace it with the queued to processing transition and an honest "waiting for a slot" message. If you stream status events, poll instead; Sume's subscribe mode is a bounded 30 second wait, not a stream.

What mistakes show up when you move between them?

Teams porting a fal client usually carry three assumptions across. Each one has a Sume-side fix, and none needs a workaround beyond the documented endpoints.

  • Assuming a queued job is stuck: on Sume queued is the normal waiting state under queue-first admission.
  • Polling in a tight loop: honor next_poll_after_seconds when present and back off otherwise.
  • Resubmitting after a client timeout: reuse the same Idempotency-Key, since a replay returns the original job instead of billing a second one.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume