A Sume job looks stuck: wait, cancel or poll the events?
Read status, then events. Queued and processing mean wait, cancel only works before generation starts, and a client timeout never cancels the job.

Do not resubmit and do not cancel on a hunch. Read GET /v1/jobs/{id}/status, then GET /v1/jobs/{id}/events, and decide from what they say: queued is a wait for a concurrency slot, processing is a wait for the generation, and a cancel works only before generation starts. A client timeout never cancels the job, so a job that you gave up on is still running, and still billing.
Read the status first
The five statuses are queued, processing, completed, failed and canceled, and the last three are terminal. The status payload is flat, with terminal, result_ready and a next_poll_after_seconds hint, so one read tells you whether to stop. If terminal is true, fetch the result when result_ready is true, and read the job for the error when it is not.
| What you see | What it means | Action |
|---|---|---|
queued | Accepted, waiting for a concurrency slot | Wait. Concurrency is a dispatch limit, so queued is normal when you are at your limit. |
processing, events show generation.submitted | Generation is underway | Wait. A cancel now returns 409 job_generation_already_started. |
processing, no generation.submitted | Started, generation not yet submitted | A cancel can still succeed. |
completed | Done | Fetch GET /v1/jobs/{id}/result. |
failed | Ended with a public error | Read the error, then decide whether to resubmit with a new key. |
| Your client timed out | Your wait ended, not the job | Read status. Never resubmit the same intent. |
Queued is capacity, not a fault
Generation concurrency is a dispatch limit, not a submit limit. Free runs one job at a time, Pro four, Startup eight and Scale twenty, and Sume accepts more valid jobs as queued while queue capacity remains: 6, 24, 48 and 120 accepted jobs in total. A queued job is therefore waiting its turn. If you keep submitting past accepted capacity, the answer is 429 queue_full on the submit, not a stuck job.
If a batch is queued behind itself, the fix is to size your waves from the generation_limits object returned on submit, which has the effective concurrency and a wave_size_hint.
Reading the events
The events route is a pull snapshot, not a stream: job.created, job.queued, job.started, generation.submitted, a terminal event, and webhook.delivery. It never shows raw provider task ids or URLs, so you can paste it into a ticket. The two events to look for are job.started, which tells you the job left the queue, and generation.submitted, which tells you that a cancel is no longer possible.
If a job has been processing for much longer than you expect for that product, include the request_id and the events in a support message. Do not create a second job to see whether it finishes sooner, because both will bill.
A useful habit is to record the status and the last event name each time you poll. When someone asks why a job took eleven minutes, the answer is then a stored sequence, such as queued for nine minutes and processing for two, instead of a guess. That sequence also tells you whether to buy more concurrency or to fix your wave size.
If many jobs look stuck at once, check the account before the jobs. A wallet that ran out fails new submits with 402, a plan at its accepted-job capacity fails them with 429 queue_full, and a rate-limited poller gets 429 rate_limited on reads, so the jobs themselves are often fine and the reader is the problem.
Canceling on purpose
Cancel with POST /v1/jobs/{id}/cancel. It succeeds only before generation work starts. Afterward the API answers 409 job_generation_already_started with details.cancelable: false, and the job runs to completion. Canceling a job that is already canceled is idempotent and returns the same canceled job.
On Formats, the run itself has a ceiling: expires_at is 90 minutes after creation, or earlier if the run goes silent, and then Sume finalizes it as failed. A run you stopped caring about can be canceled with its own cancel route while it is still cancelable, and otherwise it ends by itself.
Sources
Related posts
More in Developers
- Sume timeouts in one table: 30 s, 55 s, 10 s, 90 minutes
Every wait in the Sume API has its own number: sync 30 s, jobs_wait 55 s, webhook attempts 10 s, SDK helpers 10 and 20 minutes, Format runs 90 minutes.
- Can a Sume webhook arrive twice? Build an idempotent receiver
Sume retries failed webhook deliveries up to 10 times and Redeliver replays a real event, so one terminal event can reach you twice. Dedupe on job_id or run_id.
- 10 hooks by 10 endings: a 100-variant grid in one Sume bulk queue
A 10 by 10 hook and ending grid is exactly 100 items, the bulk queue maximum. How to build the items array, pick concurrency up to 16, and read the result.
- Test a Sume webhook receiver with signed fixtures, no paid job needed
Generate sume-v1 signatures yourself and test six cases: good, rotated, reserialized, stale, empty-secret and unknown event. Python code that runs as is.
Written by Sume