Sume job failed with worker_timeout: poll again or retry?

A Sume job error in the worker_timeout or generation_timeout category means poll status or retry later. runtime_unavailable means retry later, gently.

5 min readSume
All posts

When a Sume job error carries the category worker_timeout or generation_timeout, the documented next action is to poll status or retry later. For runtime_unavailable it is to retry later without hammering. None of these says to change your input, which separates them from validation and generation_rejected.

The practical point is what not to do: do not resubmit the paid request in a tight loop, and do not treat a timeout category as proof that nothing ran.

Which category maps to which action?

The Errors and rate limits page lists the common job error categories. Failed jobs expose public metadata such as category, stage, retryability, retry-after seconds, public reason and next action. Internal provider payloads are not public fields.

Job error categories and the documented next action (read 2026-10-02)
CategoryTypical next action
validationFix input
authCheck API key and workspace access
quotaAdd funds or lower request cost
queueRetry later with the same idempotency key
generation_unavailableRetry later
generation_rejectedInspect events and fix unsupported input
generation_timeoutPoll status or retry later
runtime_unavailableRetry later; do not retry aggressively
worker_timeoutPoll status or retry later
internalInspect events and contact support with request/job id

What is the difference between a timeout and a rejection?

A rejection is about your input. generation_rejected says to inspect events and fix unsupported input, so a retry with the same input is unlikely to help. A timeout is about time, not input: the same request can succeed later, which is why the docs point to polling or retrying.

That also explains why timeouts sit next to a rule from the jobs page: do not resubmit the original paid request just because a local process timed out. A client-side timeout does not cancel the job. It keeps running and keeps billing, and you have only stopped watching.

What does a safe response look like?

Read the failure from the job record, then branch on the category and the retry hints it carries:

  • Store the job id and request id from the submit response. Both are safe to share with support.
  • Read GET /v1/jobs/{id}/events to see the public timeline before deciding anything.
  • For generation_timeout or worker_timeout, poll status_url with backoff and honor next_poll_after_seconds when present.
  • If you do resubmit, send the same Idempotency-Key so a retry returns the original job instead of billing a second one.
  • For runtime_unavailable, back off and retry later. Do not loop quickly.

Should I read events before polling?

Yes, it is cheap and it is public. GET /v1/jobs/{id}/events returns a timeline with events such as job.created, job.queued, job.started, generation.submitted, job.completed, job.failed, job.canceled and webhook.delivery. Seeing whether the job ever reached job.started helps you decide whether this was a queue problem or a run problem.

The events endpoint is a pull snapshot, not a stream. There is no SSE or WebSocket transport on the Developer API today, so read it when you need it rather than waiting for pushes.

What do these docs not promise?

They do not say how long a timeout takes to resolve, and they do not guarantee a job will complete after a retry. They also do not describe whether a failed job's reservation is released for every category: the generation admission page says failed jobs and failed queue admission release or refund the reservation where applicable, so check the job and your balance rather than assuming.

If the category is internal, stop retrying and contact support with the request id and job id. For a different angle on the same failure family, see what happens when a sync wait times out.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume