Claude Code 2.1.288 background session fix: Sume jobs keep running

Claude Code 2.1.288 fixed background sessions ending on plugin reload. A Sume job outlives the session; how to find it and avoid paying twice.

5 min readSume
All posts

When a Claude Code background session dies mid-task, the Sume jobs it started are still running and still billing, and the recovery is to find them by listing jobs, not to ask the agent to create them again. Claude Code 2.1.288 (2026-10-02) fixed background sessions ending when a plugin was reloaded or disabled while timers or reads were running (read 2026-10-03). If you hit that bug before updating, or any session end, this is the checklist for the jobs left behind.

The Claude Code fix is in the changelog. The Sume recovery rules are in Jobs and results and MCP tools and gates.

Why the job survives

A Sume generation is a server-side job. The MCP call that created it returns a job id; the session only waits for it. When the session ends, nothing tells the job to stop. Sume documents jobs_cancel as a write tool with an idempotency_key, so cancelling is an explicit act, and wait_for: "any" on a batch wait still reports every id while the remaining jobs continue and still bill.

That is the right behavior for long media jobs, and the reason a session crash is a recovery problem rather than a data-loss problem. The id is the handle; the work is not lost.

After a session ends, read 2026-10-03
SituationWhat happened to the jobWhat to do
Session ended while waitingJob keeps runningRead jobs_status or jobs_wait on the id
Id not rememberedJob exists in the workspaceCall jobs_list and match by time and tool
Unsure it was createdMay or may not existRetry the create with the same idempotency_key
Job no longer wantedStill billingjobs_cancel with a new idempotency key

The recovery order

Follow the order below, because the first step is free and the last one is the only one that can spend.

  • jobs_list to find recent jobs, then jobs_get on candidates.
  • If the job is running, jobs_wait with the id; wait_slice_expired means retry the wait, never the create.
  • If the job finished, jobs_result (or a batch jobs_result with job_ids) returns the output.
  • If you do not know whether the create was accepted, repeat it with the original idempotency_key; a replayed key returns the original job.
  • Only if nothing matches, create anew with a new key.

Prevent the double charge

The expensive mistake is recreating a job because the session said it failed. Make the key a function of the task, not the attempt. Keep a small file next to the work that records each key and, once known, its job id, so a new session can read it before acting. A shell sketch is enough.

#!/bin/sh
# record-job.sh <task> <job_id>  (append once the id is known)
set -eu
mkdir -p .sume
printf '%s\t%s\t%s\n' "$(date -u +%FT%TZ)" "$1" "$2" >> .sume/jobs.tsv
tail -n 5 .sume/jobs.tsv

Tell the agent in its instructions to read .sume/jobs.tsv first and to call jobs_status on any id listed before creating anything. Do not put signed URLs or tokens in that file; the docs say to prefer public ids and media.sume.com URLs in reports. If the session was killed because a tool call exceeded a client timeout, that is a separate cause with the same fix: the client timeout does not cancel the job, and a held jobs_wait call is capped at 55 seconds per call, so retry the wait.

A note on cancelling

Cancel only what you are sure you do not want. jobs_cancel is a write tool and needs an idempotency_key, and a job that already finished cannot be un-spent. If you cancel a job whose result you could still use, you have to pay to make it again. Read jobs_status first, and when in doubt, let a short job finish and discard the output.

If you do cancel, record the job id as cancelled in your own log, so a later session does not read an old entry and recreate it. Cancelling twice is harmless; the second call returns the same canceled job.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume