Claude Code 25,000-token MCP output cap and Sume batch jobs_result

Claude Code truncates MCP output past 25,000 tokens by default. Sume include_results and batch jobs_result keep wave answers small. Settings and loop.

5 min readSume
All posts

Claude Code warns at 10,000 tokens of MCP output and caps it at 25,000 unless you set MAX_MCP_OUTPUT_TOKENS. A Sume wave of 20 finished jobs read back in one answer can be large, so read it in batches and let Sume say what did not fit through results_omitted.

Output limits, read 2026-10-08
SettingValue
Claude Code warning threshold10,000 tokens, fixed
Claude Code default limit25,000 tokens
OverrideMAX_MCP_OUTPUT_TOKENS env var before start
Per-tool overrideanthropic/maxResultSizeChars in tool metadata

What Sume gives you

Remote jobs_wait and jobs_result both take 1 to 20 job ids. With include_results: true, a wait returns each completed id's result in results[]. Results that do not fit one answer are named in results_omitted.job_ids, and you read those with one batch jobs_result.

Sume batch reads, read 2026-10-08
FieldMeaning
job_ids1 to 20 ids
wait_forall (default) or any
include_resultsAdds results[] for completed ids
results_omitted.job_idsIds to fetch with jobs_result
partial_failure.failed_job_idsIds needing a second read

Plan for a wave

Keep the outputs inside the client cap by splitting a large wave. Use waits that return status, then read results in groups, and check ok on each entry.

  • Submit the wave with one idempotency_key per job.
  • Call jobs_wait with job_ids and no include_results for big payloads.
  • Read finished ids with jobs_result in groups of a few.
  • An id still running returns job_not_completed; read it again later.
  • Raise MAX_MCP_OUTPUT_TOKENS only if smaller groups are not enough.

Settings

Set the variable before launching the client. It applies to tools that do not declare their own limit.

export MAX_MCP_OUTPUT_TOKENS=50000
claude

Why not just raise the cap

Raising MAX_MCP_OUTPUT_TOKENS makes each answer bigger, and everything the model reads costs context. A wave of 20 results that fits in the cap still fills the conversation. The Sume batch fields let you keep the model's context for decisions and keep bulk data in public ids and URLs.

Reading results by group

A practical rule is to read results in the order the agent needs them. Ask for the first group, act on it, then ask for the next. Because jobs_result returns results[] in request order with an ok flag per entry, the agent can skip failures and name them to you. Sume public ids and media.sume.com URLs are preferred in reports over signed URLs, which should not be pasted into logs.

Also remember that the 10,000-token warning is fixed and cannot be changed. If you see it often, the agent is probably reading full results where a status would do. Ask it to call jobs_status for progress checks and to keep jobs_result for the end. For crawl or inspect outputs, which tend to be large, ask for one id at a time and have the agent write only the fields it needs into its notes. That keeps both the client cap and the model context in good shape while the wave runs. Sume's per-answer budget is a separate limit on the server side, so the client setting alone does not remove it; batching reads is the reliable fix on both sides.

Before you rely on this setup, run a short acceptance test with a read-only credential. Connect, call mcp_health, call tools_list, and read one job with jobs_status. Record the tool count you see, so you can notice later if a credential change alters it. Then repeat the test after any config edit. A five-minute test like this catches most wiring mistakes before they cost money, and it gives you a baseline to compare against when something behaves differently next week.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume