Claude Code CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS and Sume jobs_wait

Claude Code can move a long MCP call to the background after a threshold. Sume's jobs_wait returns within 55 seconds, so set the threshold above that.

5 min readSume
All posts

Leave CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS above 55000 and Claude Code will not background a Sume jobs_wait call, because Sume never holds one longer than 55 seconds. If you set it lower, the wait may be moved to the background; the Sume job is unaffected either way, since it runs on Sume's side.

The variable is described in Claude Code's MCP docs (read 2026-10-02) as the threshold, in milliseconds, for automatically backgrounding long tool calls; the page gives a default of 120000 (2 minutes), and 0 disables it; it requires Claude Code v2.1.212 or later. A companion variable, CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1, turns off automatic backgrounding of long tool calls. Sume's limits come from Jobs and results.

What does auto-backgrounding do to an MCP call?

When a tool call runs past the threshold, Claude Code moves it to a background task so the session can keep going. The changelog for 2.1.283 notes a related fix: progress notifications were being discarded once a long-running call moved to the background, and the background task now shows the latest progress (Claude Code changelog, read 2026-10-02).

The default threshold is 120000 ms (2 minutes), already above Sume's 55 second slice. What you can control is the setting itself: raise it, lower it, or set it to 0.

How long can one Sume call actually run?

Most Sume tools answer quickly because paid generation is asynchronous: generate_video returns a job, and you read it later. The only call built to sit open is jobs_wait. On remote MCP its timeout_seconds defaults to 50 and is capped at 55; larger values up to 600 are accepted and clamped, and the response reports wait_slice_clamped. The docs explain why: a wait is one HTTP request held open with nothing transferring, and every edge closes such a request eventually.

So a single Sume call that exceeds a threshold of 60 seconds or more should not happen. Waiting for a ten-minute render means repeating the wait on the same ids, not asking for a longer one.

Threshold against a 55-second jobs_wait slice, read 2026-10-02
Threshold settingSume jobs_wait (up to 55 s)What to expect
0 (disabled)Runs in the foregroundClaude waits for each slice
120000 (the default)Finishes before the thresholdNever backgrounded
30000Can pass the threshold on a long sliceMay be backgrounded; job keeps running
Backgrounding off via env varRuns in the foregroundSame as disabled

What if a wait does get backgrounded?

Nothing is lost on Sume's side. The job is billed and processed whether or not Claude is watching. When the background task reports wait_slice_expired, or you return to the session, call jobs_wait again with the same job_id or job_ids. Do not resubmit the paid create: the docs say never to, and a new idempotency_key would start a second job. Reusing the same key on a retry is the safe pattern for the create itself.

For a batch, prefer one jobs_wait with up to 20 job_ids and include_results: true over several single waits. One backgrounded task is easier to follow than six.

What should I set?

If you run mostly Sume generation through Claude Code, the simplest configuration is to leave the default alone and keep your own per-server timeout in .mcp.json out of the way. If you want fewer surprises, set the threshold to a value above 55000, such as 60000, so a normal wait stays in the foreground while a genuinely stuck call can still move aside.

Related reading: Claude Code MCP tool timeout covers the per-server timeout value, and progress notifications on a background call covers the 2.1.283 fix. Sume does not stream progress over MCP; it exposes bounded waits and jobs_status reads.

export CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS=60000
claude

Does this change what Sume bills?

No. Billing follows the job, not the client. Sume reserves cost at submit for a paid create and settles when the job finishes, so whether Claude Code keeps a call in the foreground, backgrounds it, or aborts it has no effect on the amount. That is why the safe habits are the same in every case: submit once with a stable idempotency_key, run a dry_run or generation_admission_preview first when the spend is large, and read usage_get and balance_get afterwards to see what happened.

If you want a belt-and-braces setup for unattended sessions, pass max_spend_usd on paid calls. The docs say it is enforced only when you provide it, so it is opt-in per call. Pair that with a permission rule in Claude Code so a paid tool never runs without a prompt, as described in asking before paid MCP tools.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume