Resume a video chain after a worker crash from stored job ids

Your worker died between generate and trim. Read the stored job id from /v1/jobs, do not resubmit paid work, and continue from the first step with no result.

4 min readSume
All posts

If a worker crashes in the middle of a generate, trim and captions chain, do not submit the step again. Read the stored job id with GET /v1/jobs/:id/status, and continue from the first step that has no completed result. Sume jobs are durable: a client-side timeout or a dead process does not cancel the job, and it keeps running and billing, so the job id is the thing you must have saved.

What to persist at each step

The docs' first instruction is to store the job id from submit responses so you can recover work after restarts. Persist one row per step, written as soon as the submit returns, before you start waiting.

Per-step record that makes a chain resumable (read 2026-10-05)
FieldSourceWhy you need it
stepYour chain definitionWhich of generate, trim or captions this row is
idempotency_keyYou choose itA retry of the submit returns the original job
job_idrequest_id in the submit responseThe handle for status, result, events and cancel
result_urlThe step's result, once result_readyThe input for the next step

The restart procedure

Walk the rows in order. For each row with a job_id, read GET /v1/jobs/:id/status and branch on terminal: if it is false, keep polling; if it is true and sume_status is completed, read the result and move on; if it is failed or canceled, read the job record for its public error and decide whether a new submit is justified. For a row with no job_id, submit with the stored idempotency_key: if the first request did create a job before the crash, you get that job back.

curl https://api.sume.com/v1/jobs/job_123/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/job_123/result \
  -H "Authorization: Bearer $SUME_API_KEY"

Three traps

The first is reading /result on a job that is not completed. It answers 409 job_not_completed and never an empty result, so poll status first and fetch only when result_ready is true. The second is losing the key: a fresh Idempotency-Key on a retry creates a second paid job. The third is the job-ownership rule: an API key reads only the jobs its own member created, so the restarted worker must use a key from the same member, or the read is a 404 not_found.

If you cannot find a job id at all, GET /v1/jobs lists jobs for the key's member, which is the documented way to look for work after a restart.

A restart routine in four moves

On startup, load every chain row that is not marked done. For each one, look at the last stored step. If the step has a job id, read GET /v1/jobs/:id/status and branch on terminal. If it is terminal and result_ready, read the result and move on to the next step. If it is still running, keep polling with next_poll_after_seconds.

If a row has a step name and key but no job id, the crash happened between the submit and the write. Resend the same body with the same key. The API returns the existing job if one was created, so you cannot double-submit and double-pay.

If the status reads failed or canceled, record it and stop the chain; the next step has no valid input. Do not guess a replacement.

  • Load unfinished rows.
  • Poll stored job ids first.
  • Resend with the stored key when no id was saved.
  • Mark the chain failed on failed or canceled.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume