Claude Code RETRY_WATCHDOG gives up after 3 timeouts: Sume jobs

Unattended Claude Code sessions using CLAUDE_CODE_RETRY_WATCHDOG now stop after three timeouts. Persist Sume job ids so a restart resumes, not resubmits.

5 min readSume
All posts

The Claude Code 2.1.288 release notes in the changelog, read 2026-10-03, change how unattended sessions behave. Sessions using CLAUDE_CODE_RETRY_WATCHDOG previously retried for hours after a failed long stream; now they stream again and give up after three timeouts. The same release says non-interactive sessions and subagents continue from the partial response after a mid-response API timeout, and that MCP calls could run twice on a result over 16 MB or an unparseable one.

What this means for a Sume run

An unattended session that submits a Sume job and then waits can now end earlier than before. That is good for cost and bad if you assumed the session would wait it out. The job does not stop when the session does: the Sume jobs guide says a client timeout does not cancel the job and it keeps billing.

2.1.288 behaviours that touch Sume runs, read 2026-10-03.
ChangeEffect on a Sume run
Gives up after three timeoutsSession may end while a job runs
Continues from a partial responseDo not assume the create call was repeated or skipped; check the job
MCP calls could run twice on a very large or unparseable resultSame reason to use an idempotency_key
Re-authenticate prompt for more OAuth scopeUnattended runs need the scope up front

Make the session restartable

Write the job id to a file the moment the create tool returns it, along with the idempotency key. On a restart, read the file first. If it holds a job id, call jobs_wait or jobs_result instead of creating again. If a crash happened before the job id was written, repeating the create with the same idempotency_key returns the same job rather than a second one.

The Claude Code MCP docs list related limits: CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT is 5 minutes for HTTP, and MAX_MCP_OUTPUT_TOKENS is 25,000. Keep result envelopes small with include_results only when you need them.

Checklist

  • Choose the idempotency key before the session starts, for example from the batch id.
  • Cap spend with max_spend_usd.
  • Request mcp:write up front if the run must create, since an unattended session cannot answer a re-authenticate prompt.
  • Reuse the same ids in jobs_wait after wait_slice_expired.
  • Check balance and job state at the start of every resumed run.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume