Resume a video chain after a worker crash from stored job ids
Your worker died between generate and trim. Read the stored job id from /v1/jobs, do not resubmit paid work, and continue from the first step with no result.

If a worker crashes in the middle of a generate, trim and captions chain, do not submit the step again. Read the stored job id with GET /v1/jobs/:id/status, and continue from the first step that has no completed result. Sume jobs are durable: a client-side timeout or a dead process does not cancel the job, and it keeps running and billing, so the job id is the thing you must have saved.
What to persist at each step
The docs' first instruction is to store the job id from submit responses so you can recover work after restarts. Persist one row per step, written as soon as the submit returns, before you start waiting.
| Field | Source | Why you need it |
|---|---|---|
step | Your chain definition | Which of generate, trim or captions this row is |
idempotency_key | You choose it | A retry of the submit returns the original job |
job_id | request_id in the submit response | The handle for status, result, events and cancel |
result_url | The step's result, once result_ready | The input for the next step |
The restart procedure
Walk the rows in order. For each row with a job_id, read GET /v1/jobs/:id/status and branch on terminal: if it is false, keep polling; if it is true and sume_status is completed, read the result and move on; if it is failed or canceled, read the job record for its public error and decide whether a new submit is justified. For a row with no job_id, submit with the stored idempotency_key: if the first request did create a job before the crash, you get that job back.
curl https://api.sume.com/v1/jobs/job_123/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/job_123/result \
-H "Authorization: Bearer $SUME_API_KEY"Three traps
The first is reading /result on a job that is not completed. It answers 409 job_not_completed and never an empty result, so poll status first and fetch only when result_ready is true. The second is losing the key: a fresh Idempotency-Key on a retry creates a second paid job. The third is the job-ownership rule: an API key reads only the jobs its own member created, so the restarted worker must use a key from the same member, or the read is a 404 not_found.
If you cannot find a job id at all, GET /v1/jobs lists jobs for the key's member, which is the documented way to look for work after a restart.
A restart routine in four moves
On startup, load every chain row that is not marked done. For each one, look at the last stored step. If the step has a job id, read GET /v1/jobs/:id/status and branch on terminal. If it is terminal and result_ready, read the result and move on to the next step. If it is still running, keep polling with next_poll_after_seconds.
If a row has a step name and key but no job id, the crash happened between the submit and the write. Resend the same body with the same key. The API returns the existing job if one was created, so you cannot double-submit and double-pay.
If the status reads failed or canceled, record it and stop the chain; the next step has no valid input. Do not guess a replacement.
- Load unfinished rows.
- Poll stored job ids first.
- Resend with the stored key when no id was saved.
- Mark the chain failed on
failedorcanceled.
Sources
Related posts
More in Developers
- Retry a Kling motion control submit without paying twice
All four Sume motion-control routes take an Idempotency-Key. Resend the same body and key after a timeout and you get the original job, not a second reserve.
- Retry a Sume video submit safely: Idempotency-Key and the 409 conflict
Send the same Idempotency-Key and the same body to retry a Sume video submit without a second charge. A different body under the same key returns 409.
- Retry a video-trim: same key for the same body, new key for new start
A trim that timed out can be retried with the same Idempotency-Key and body. Changing start, duration or precision is a new operation and needs a new key.
- Rolling a webhook secret: Stripe's 24-hour overlap vs Sume's header
Stripe can keep an old signing secret live for up to 24 hours. Sume sends one sume-v1 entry per live secret; verify any match. Python verifier included.
Written by Sume