LangGraph request_drain and GraphDrained: resume a Sume job later
LangGraph 1.2 can drain a run after the current superstep. If a Sume job is in flight, save its id so the resumed run polls instead of resubmitting.

If a LangGraph run is drained while a Sume job is in flight, the job keeps running; your state must carry the job id so the resumed run polls it instead of submitting again. Sume jobs are durable records, so a drain and a deploy do not lose the work, only your handle on it.
The LangChain Python changelog (read 2026-10-02) describes graceful shutdown in LangGraph v1.2.0: create a RunControl, call request_drain(), and the run stops cooperatively after the current superstep, raising GraphDrained so it can be resumed later.
What does a drain do to a Sume job?
Nothing. A drain stops your graph between supersteps; it does not talk to Sume. A job already submitted continues through queued, processing, and a terminal state, and it bills while you are drained (Jobs and results). Sume says to store the job id from submit responses so an integration can recover work after process restarts, which is the drain case exactly.
The risk is the superstep boundary. If the submit and the state write that saves the job id land in different supersteps, a drain between them leaves a paid job you cannot find.
How do I make submit and save one step?
Do both in the same node and write the job id as that node's return value, so the checkpoint that follows already holds it. Use a stable Idempotency-Key built from the business item. If the drain still falls in the gap, the resumed run repeats the submit with the same key and gets the original job back, with no second charge.
| Drain lands | State holds | Resume does |
|---|---|---|
| Before submit | No job id | Submit with the item key |
| After submit, before checkpoint | No job id | Submit again with the same key, get the same job |
| After checkpoint, job running | Job id | Poll status until terminal |
| After terminal, before fetch | Job id, terminal | Fetch the result once |
How should the resumed run wait?
Poll GET /v1/jobs/{id}/status with backoff, honouring next_poll_after_seconds when present, and stop on terminal: true. Fetch GET /v1/jobs/{id}/result only when result_ready is true; before that it answers 409 job_not_completed.
If the graph should not hold a worker while a video renders, ask for a webhook instead: submit with mode: "webhook" and a public HTTPS webhook_url, and let the callback resume the graph. Webhooks are terminal-only (job.completed, job.failed, job.canceled) and signed, and Sume tells you to keep polling available as a backup (Webhooks).
What should I check before a deploy that drains runs?
List the runs that hold a non-terminal job id. For each, confirm the checkpoint contains the id and the item key. Sume's GET /v1/jobs lists jobs the key's member created, so you can cross-check that every in-flight job has a graph run that knows about it.
After the deploy, resume the drained runs and watch for submits. A correct resume of a run with a saved id should send zero submit requests and only status reads. If the resumed run submits, either the id never reached state or the node ignores it; fix that before the next drain.
Remember that spend continues while you are drained. A drain that lasts an hour does not pause a ten-minute render. It means the job completes, bills, and waits for you, which is fine as long as the result is fetched later from the saved id.
What about a drain during a long jobs_wait?
If your graph calls Sume over remote MCP, one jobs_wait holds at most 55 seconds (default 50) and answers wait_slice_expired when the slice ends; repeat the wait with the same ids and never resubmit the paid create. A drain between slices is harmless for the same reason. Sume does not resume your graph for you and has no push stream, so the resume trigger is yours: a scheduler, a webhook, or an operator. Before you ship, read the live contract for every route you call at https://api.sume.com/reference/json, which Sume's docs name as the schema source of truth, and re-read the linked docs pages: limits, scopes and error codes change faster than blog posts do. Treat any number in this post as a snapshot dated 2026-10-02, and prefer the effective fields your own responses return, such as generation_limits, over a static table.
Sources
Related posts
More in Developers
- LinkedIn API version 202510 sunsets Oct 15, 2026: what to change
LinkedIn Marketing version 202510 is sunset on October 15, 2026. Pin Linkedin-Version 202609 and add a check so a stale pin fails in CI, not in production.
- LinkedIn Videos API: 4 MB parts, ETags and finalizeUpload
LinkedIn video upload is initializeUpload, PUT each 4 MB part, collect the ETags, then finalizeUpload. A part splitter and the order rules that break uploads.
- MAI-Image-2.6 allows 6 requests a minute at tier 1; Sume queues
Foundry rates MAI-Image-2.6 at 6 RPM on tier 1 and 429s past it. Sume accepts valid image jobs as queued until a plan slot opens. Compare the two behaviours.
- MAI-Image-2.6 returns base64 PNG; Sume returns a hosted image URL
Foundry returns MAI-Image-2.6 as b64_json PNG only. Sume returns a signed media URL in data[].url, in png, jpeg or webp per model. How to handle each in code.
Written by Sume