fal queue_position and logs vs Sume job status, events, usage
fal returns queue_position, runner logs and inference_time. Sume gives status, an events timeline and per-job usage. Which one tells you why a job is slow?

What does fal tell you while a request runs?
Three things beyond the state name. fal's queue documentation says status() takes an optional with_logs=True (Python) or logs: true (JavaScript) to return runner output, includes queue_position while a request is queued, and includes metrics.inference_time once complete. Sume's job status gives you the state, a timeline of events, and the cost of the job, but not runner logs.
What does Sume give you for the same questions?
For "where is my job", GET /v1/jobs/:id/status returns the lightweight state with terminal, result_ready and next_poll_after_seconds, plus a queue-shaped status that maps one-to-one onto sume_status. The jobs docs do not describe a per-job queue position on that payload.
For "what happened", GET /v1/jobs/:id/events returns a public timeline: job.created, job.queued, job.started, generation.submitted, job.completed, job.failed, job.canceled and webhook.delivery. Public events do not expose raw provider task ids or URLs. For "how much did it cost", GET /v1/usage?job_id=... sums the ledger rows for that job. See Usage.
Side by side
Observability fields, from each vendor's own docs.
| Question | fal | Sume |
|---|---|---|
| Is it waiting? | IN_QUEUE, with queue_position | queued; plan queue counts in generation_limits at submit |
| What is it doing? | Runner logs via with_logs / logs: true | Event timeline; no runner logs |
| How long did inference take? | metrics.inference_time when complete | Compare the job.started and job.completed events |
| What did it cost? | Not on the queue page read | /v1/usage?job_id= with debited_usd |
| Pushed progress | Status stream over SSE or polling | Terminal webhooks only; poll or read events |
How do I find out why a Sume job is slow?
Start with the queue, because most slow Sume jobs are waiting rather than running. The submit response carries generation_limits with active_generation_jobs, queued_generation_jobs and queue_capacity_remaining. If your workspace is at its processing limit, new jobs sit in queued by design. Then read the events to see when job.started landed.
# job id comes from the submit response
curl https://api.sume.com/v1/jobs/job_123/events \
-H "Authorization: Bearer $SUME_API_KEY"
curl "https://api.sume.com/v1/usage?job_id=job_123" \
-H "Authorization: Bearer $SUME_API_KEY"What do you lose by moving from fal to Sume?
If you tail runner logs to debug a model, you lose that: Sume does not publish them. If you poll queue_position to drive a progress bar, replace it with the queued to processing transition and an honest "waiting for a slot" message. If you stream status events, poll instead; Sume's subscribe mode is a bounded 30 second wait, not a stream.
What mistakes show up when you move between them?
Teams porting a fal client usually carry three assumptions across. Each one has a Sume-side fix, and none needs a workaround beyond the documented endpoints.
- Assuming a queued job is stuck: on Sume
queuedis the normal waiting state under queue-first admission. - Polling in a tight loop: honor
next_poll_after_secondswhen present and back off otherwise. - Resubmitting after a client timeout: reuse the same
Idempotency-Key, since a replay returns the original job instead of billing a second one.
Sources
Related posts
More in Comparisons
- fal webhook redirect 3xx is never retried: what Sume does instead
fal treats a 3xx from your webhook URL as a permanent failure. Sume does not follow redirects either, but counts a 3xx as a failed attempt, not an end.
- fal webhook ED25519 and JWKS vs Sume's HMAC-SHA256 check
fal signs webhooks with ED25519 keys fetched from a JWKS URL. Sume signs HMAC-SHA256 over timestamp.body with a workspace secret. The two checks, side by side.
- fal webhook retries: 31 attempts, 15 s timeout, vs Sume's 10
fal retries a failed webhook with backoff up to 31 times while the stored result lasts; Sume makes up to 10 attempts with a 10 second timeout. What to build.
- Fastest AI image model API: what the October 2026 claims say
Flare says half the latency of GPT Image 2, MAI-Image-2.6-Flash says 2.8x faster than GPT-Image-2-Medium. None are comparable. A timing script for Sume models.
Written by Sume