LangGraph error_handler after retries: do not resubmit a Sume job
LangGraph 1.2 node error handlers run after retries are exhausted. For a Sume call, the handler should read job status, not submit a second paid request.

In LangGraph 1.2, a node-level error_handler runs after retries are exhausted, which is the right place to ask Sume what happened to a job, and the wrong place to submit it again. Keep the job id in state, read GET /v1/jobs/{id}/status, and only treat a Sume failed status as failed.
The LangChain Python changelog (read 2026-10-02) lists, for v1.2.0 of May 12, 2026, an error_handler= parameter on add_node() that runs a recovery function after retry exhaustion and receives a typed NodeError. It says timeouts and error handlers are Python-only.
Why is retry-then-handler risky around a paid call?
A retry policy cannot tell whether the first request reached Sume. If a node submits a video, the connection drops after Sume accepted it, and the retry submits again without a key, you now own two paid jobs. Sume's docs are blunt on this: do not resubmit a paid request just because a local process timed out (Jobs and results).
The fix is the Idempotency-Key header. Retrying the submit itself is fine when you reuse the same key: the retry returns the original job. The key is for exact retries only; Sume answers 409 idempotency_conflict when it is reused for a different operation or payload.
Which Sume errors should the handler treat differently?
Sort errors by whether a retry can help. Generation admission gives a client-behavior column for each.
| Status and code | Retry? | Handler action |
|---|---|---|
| 400 invalid_request | No | Fix the request, route to a human or a repair node |
| 402 insufficient_credits | No | Stop the run and report balance |
| 409 idempotency_conflict | No | Key reused for a different payload; derive keys from the item |
| 429 queue_full | Later | Wait, retry with the same key |
| 429 rate_limited | Later | Back off per retry-after |
| 503 provider_capacity_exceeded | Later | Retry later with the same key unless the error says not to |
What should the handler do when the job id exists?
If state already holds a job id, the handler should ask Sume rather than act on its own guess. Read the status, and branch on terminal and sume_status: completed goes to the result fetch, failed reads the public error from the job record, canceled ends the item, and anything else is still in flight, so return a pending marker for a later node.
If state holds no job id, the submit never returned one. Retrying with the same idempotency key is then safe, because a key that matches returns the original job when one was created and creates one when none was.
- Never mint a new key inside the handler.
- Do not call
/cancelas cleanup; it is refused with409 job_generation_already_startedonce generation begins. - Log the job id with the
NodeErrorso a person can look it up.
How do I test the handler path?
Force each failure on purpose. Point a test at an invalid body to get a 400 invalid_request, a balance of zero to get 402 insufficient_credits, and a repeated key with a changed prompt to get 409 idempotency_conflict. For each one, assert the handler made no second submit and left a readable record on the NodeError path.
Then simulate a lost response: send the submit, drop the reply in your test client, and let the retry run with the same key. You should see one job on Sume's side. This is the case the idempotency key exists for, and it is the one a retry policy alone cannot cover.
Keep the handler small. It should read status, write one of a few outcomes into state, and return. Business decisions, such as whether to rerun a failed item with a changed prompt, belong in a later node where a person or a policy can see the cost.
Does this replace a human approval step?
No. An idempotency key stops duplicates; it is not approval, which Sume's MCP docs say in as many words. If a person should approve spend before the first submit, put an interrupt in front of the node, as covered in LangGraph interrupt before a paid API call. The error handler only covers what happens after approval and submit.
Sources
Related posts
More in Developers
- LangGraph request_drain and GraphDrained: resume a Sume job later
LangGraph 1.2 can drain a run after the current superstep. If a Sume job is in flight, save its id so the resumed run polls instead of resubmitting.
- LinkedIn API version 202510 sunsets Oct 15, 2026: what to change
LinkedIn Marketing version 202510 is sunset on October 15, 2026. Pin Linkedin-Version 202609 and add a check so a stale pin fails in CI, not in production.
- LinkedIn Videos API: 4 MB parts, ETags and finalizeUpload
LinkedIn video upload is initializeUpload, PUT each 4 MB part, collect the ETags, then finalizeUpload. A part splitter and the order rules that break uploads.
- MAI-Image-2.6 allows 6 requests a minute at tier 1; Sume queues
Foundry rates MAI-Image-2.6 at 6 RPM on tier 1 and 429s past it. Sume accepts valid image jobs as queued until a plan slot opens. Compare the two behaviours.
Written by Sume