LangGraph error_handler after retries: do not resubmit a Sume job

LangGraph 1.2 node error handlers run after retries are exhausted. For a Sume call, the handler should read job status, not submit a second paid request.

5 min readSume
All posts

In LangGraph 1.2, a node-level error_handler runs after retries are exhausted, which is the right place to ask Sume what happened to a job, and the wrong place to submit it again. Keep the job id in state, read GET /v1/jobs/{id}/status, and only treat a Sume failed status as failed.

The LangChain Python changelog (read 2026-10-02) lists, for v1.2.0 of May 12, 2026, an error_handler= parameter on add_node() that runs a recovery function after retry exhaustion and receives a typed NodeError. It says timeouts and error handlers are Python-only.

Why is retry-then-handler risky around a paid call?

A retry policy cannot tell whether the first request reached Sume. If a node submits a video, the connection drops after Sume accepted it, and the retry submits again without a key, you now own two paid jobs. Sume's docs are blunt on this: do not resubmit a paid request just because a local process timed out (Jobs and results).

The fix is the Idempotency-Key header. Retrying the submit itself is fine when you reuse the same key: the retry returns the original job. The key is for exact retries only; Sume answers 409 idempotency_conflict when it is reused for a different operation or payload.

Which Sume errors should the handler treat differently?

Sort errors by whether a retry can help. Generation admission gives a client-behavior column for each.

Sume submit errors and handler behavior (read 2026-10-02, from Generation admission)
Status and codeRetry?Handler action
400 invalid_requestNoFix the request, route to a human or a repair node
402 insufficient_creditsNoStop the run and report balance
409 idempotency_conflictNoKey reused for a different payload; derive keys from the item
429 queue_fullLaterWait, retry with the same key
429 rate_limitedLaterBack off per retry-after
503 provider_capacity_exceededLaterRetry later with the same key unless the error says not to

What should the handler do when the job id exists?

If state already holds a job id, the handler should ask Sume rather than act on its own guess. Read the status, and branch on terminal and sume_status: completed goes to the result fetch, failed reads the public error from the job record, canceled ends the item, and anything else is still in flight, so return a pending marker for a later node.

If state holds no job id, the submit never returned one. Retrying with the same idempotency key is then safe, because a key that matches returns the original job when one was created and creates one when none was.

  • Never mint a new key inside the handler.
  • Do not call /cancel as cleanup; it is refused with 409 job_generation_already_started once generation begins.
  • Log the job id with the NodeError so a person can look it up.

How do I test the handler path?

Force each failure on purpose. Point a test at an invalid body to get a 400 invalid_request, a balance of zero to get 402 insufficient_credits, and a repeated key with a changed prompt to get 409 idempotency_conflict. For each one, assert the handler made no second submit and left a readable record on the NodeError path.

Then simulate a lost response: send the submit, drop the reply in your test client, and let the retry run with the same key. You should see one job on Sume's side. This is the case the idempotency key exists for, and it is the one a retry policy alone cannot cover.

Keep the handler small. It should read status, write one of a few outcomes into state, and return. Business decisions, such as whether to rerun a failed item with a changed prompt, belong in a later node where a person or a policy can see the cost.

Does this replace a human approval step?

No. An idempotency key stops duplicates; it is not approval, which Sume's MCP docs say in as many words. If a person should approve spend before the first submit, put an interrupt in front of the node, as covered in LangGraph interrupt before a paid API call. The error handler only covers what happens after approval and submit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume