LangGraph request_drain and GraphDrained: resume a Sume job later

LangGraph 1.2 can drain a run after the current superstep. If a Sume job is in flight, save its id so the resumed run polls instead of resubmitting.

5 min readSume
All posts

If a LangGraph run is drained while a Sume job is in flight, the job keeps running; your state must carry the job id so the resumed run polls it instead of submitting again. Sume jobs are durable records, so a drain and a deploy do not lose the work, only your handle on it.

The LangChain Python changelog (read 2026-10-02) describes graceful shutdown in LangGraph v1.2.0: create a RunControl, call request_drain(), and the run stops cooperatively after the current superstep, raising GraphDrained so it can be resumed later.

What does a drain do to a Sume job?

Nothing. A drain stops your graph between supersteps; it does not talk to Sume. A job already submitted continues through queued, processing, and a terminal state, and it bills while you are drained (Jobs and results). Sume says to store the job id from submit responses so an integration can recover work after process restarts, which is the drain case exactly.

The risk is the superstep boundary. If the submit and the state write that saves the job id land in different supersteps, a drain between them leaves a paid job you cannot find.

How do I make submit and save one step?

Do both in the same node and write the job id as that node's return value, so the checkpoint that follows already holds it. Use a stable Idempotency-Key built from the business item. If the drain still falls in the gap, the resumed run repeats the submit with the same key and gets the original job back, with no second charge.

Where a drain can land (read 2026-10-02; LangGraph terms from the changelog)
Drain landsState holdsResume does
Before submitNo job idSubmit with the item key
After submit, before checkpointNo job idSubmit again with the same key, get the same job
After checkpoint, job runningJob idPoll status until terminal
After terminal, before fetchJob id, terminalFetch the result once

How should the resumed run wait?

Poll GET /v1/jobs/{id}/status with backoff, honouring next_poll_after_seconds when present, and stop on terminal: true. Fetch GET /v1/jobs/{id}/result only when result_ready is true; before that it answers 409 job_not_completed.

If the graph should not hold a worker while a video renders, ask for a webhook instead: submit with mode: "webhook" and a public HTTPS webhook_url, and let the callback resume the graph. Webhooks are terminal-only (job.completed, job.failed, job.canceled) and signed, and Sume tells you to keep polling available as a backup (Webhooks).

What should I check before a deploy that drains runs?

List the runs that hold a non-terminal job id. For each, confirm the checkpoint contains the id and the item key. Sume's GET /v1/jobs lists jobs the key's member created, so you can cross-check that every in-flight job has a graph run that knows about it.

After the deploy, resume the drained runs and watch for submits. A correct resume of a run with a saved id should send zero submit requests and only status reads. If the resumed run submits, either the id never reached state or the node ignores it; fix that before the next drain.

Remember that spend continues while you are drained. A drain that lasts an hour does not pause a ten-minute render. It means the job completes, bills, and waits for you, which is fine as long as the result is fetched later from the saved id.

What about a drain during a long jobs_wait?

If your graph calls Sume over remote MCP, one jobs_wait holds at most 55 seconds (default 50) and answers wait_slice_expired when the slice ends; repeat the wait with the same ids and never resubmit the paid create. A drain between slices is harmless for the same reason. Sume does not resume your graph for you and has no push stream, so the resume trigger is yours: a scheduler, a webhook, or an operator. Before you ship, read the live contract for every route you call at https://api.sume.com/reference/json, which Sume's docs name as the schema source of truth, and re-read the linked docs pages: limits, scopes and error codes change faster than blog posts do. Treat any number in this post as a snapshot dated 2026-10-02, and prefer the effective fields your own responses return, such as generation_limits, over a static table.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume