Claude Code CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS and Sume jobs_wait
Claude Code can move a long MCP call to the background after a threshold. Sume's jobs_wait returns within 55 seconds, so set the threshold above that.

Leave CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS above 55000 and Claude Code will not background a Sume jobs_wait call, because Sume never holds one longer than 55 seconds. If you set it lower, the wait may be moved to the background; the Sume job is unaffected either way, since it runs on Sume's side.
The variable is described in Claude Code's MCP docs (read 2026-10-02) as the threshold, in milliseconds, for automatically backgrounding long tool calls; the page gives a default of 120000 (2 minutes), and 0 disables it; it requires Claude Code v2.1.212 or later. A companion variable, CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1, turns off automatic backgrounding of long tool calls. Sume's limits come from Jobs and results.
What does auto-backgrounding do to an MCP call?
When a tool call runs past the threshold, Claude Code moves it to a background task so the session can keep going. The changelog for 2.1.283 notes a related fix: progress notifications were being discarded once a long-running call moved to the background, and the background task now shows the latest progress (Claude Code changelog, read 2026-10-02).
The default threshold is 120000 ms (2 minutes), already above Sume's 55 second slice. What you can control is the setting itself: raise it, lower it, or set it to 0.
How long can one Sume call actually run?
Most Sume tools answer quickly because paid generation is asynchronous: generate_video returns a job, and you read it later. The only call built to sit open is jobs_wait. On remote MCP its timeout_seconds defaults to 50 and is capped at 55; larger values up to 600 are accepted and clamped, and the response reports wait_slice_clamped. The docs explain why: a wait is one HTTP request held open with nothing transferring, and every edge closes such a request eventually.
So a single Sume call that exceeds a threshold of 60 seconds or more should not happen. Waiting for a ten-minute render means repeating the wait on the same ids, not asking for a longer one.
| Threshold setting | Sume jobs_wait (up to 55 s) | What to expect |
|---|---|---|
| 0 (disabled) | Runs in the foreground | Claude waits for each slice |
| 120000 (the default) | Finishes before the threshold | Never backgrounded |
| 30000 | Can pass the threshold on a long slice | May be backgrounded; job keeps running |
| Backgrounding off via env var | Runs in the foreground | Same as disabled |
What if a wait does get backgrounded?
Nothing is lost on Sume's side. The job is billed and processed whether or not Claude is watching. When the background task reports wait_slice_expired, or you return to the session, call jobs_wait again with the same job_id or job_ids. Do not resubmit the paid create: the docs say never to, and a new idempotency_key would start a second job. Reusing the same key on a retry is the safe pattern for the create itself.
For a batch, prefer one jobs_wait with up to 20 job_ids and include_results: true over several single waits. One backgrounded task is easier to follow than six.
What should I set?
If you run mostly Sume generation through Claude Code, the simplest configuration is to leave the default alone and keep your own per-server timeout in .mcp.json out of the way. If you want fewer surprises, set the threshold to a value above 55000, such as 60000, so a normal wait stays in the foreground while a genuinely stuck call can still move aside.
Related reading: Claude Code MCP tool timeout covers the per-server timeout value, and progress notifications on a background call covers the 2.1.283 fix. Sume does not stream progress over MCP; it exposes bounded waits and jobs_status reads.
export CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS=60000
claudeDoes this change what Sume bills?
No. Billing follows the job, not the client. Sume reserves cost at submit for a paid create and settles when the job finishes, so whether Claude Code keeps a call in the foreground, backgrounds it, or aborts it has no effect on the amount. That is why the safe habits are the same in every case: submit once with a stable idempotency_key, run a dry_run or generation_admission_preview first when the spend is large, and read usage_get and balance_get afterwards to see what happened.
If you want a belt-and-braces setup for unattended sessions, pass max_spend_usd on paid calls. The docs say it is enforced only when you provide it, so it is opt-in per call. Pair that with a permission rule in Claude Code so a paid tool never runs without a prompt, as described in asking before paid MCP tools.
Sources
Related posts
More in Integrations
- Claude Code MCP whitespace warning: a pasted Sume key with a newline
Claude Code warns when an MCP header or url has leading or trailing whitespace, often a pasted token with a newline. It does not trim it. Fix a Sume entry.
- Claude Code .mcp.json: why ${ANTHROPIC_API_KEY} reads empty for Sume
Claude Code reads credential variables like ANTHROPIC_API_KEY and NPM_TOKEN as empty in a remote url or headers. Name your Sume key variable SUME_API_KEY.
- claude mcp list and get: check the Sume server before you prompt
claude mcp list shows each server's health; claude mcp get shows one entry. Use both, then call Sume's mcp_health, to prove the connection works.
- Claude Code mcp reset-project-choices: bring back the Sume prompt
Declined or approved a project .mcp.json server by mistake? claude mcp reset-project-choices clears those choices so Claude Code asks again about Sume.
Written by Sume