Claude per-message effort: drop to low while a Sume job runs

Claude's per-message effort beta keeps the prompt cache when you change effort mid-conversation. How to use it around Sume jobs, and its Haiku 5.5 limit.

6 min readSume
All posts

To change effort partway through a conversation without throwing away the prompt cache, send a system message with empty content and an output_config.effort value, and add the beta header mid-conversation-output-config-2026-07-01. Anthropic lists Claude Haiku 5.5 among the supported models on the Claude API and Google Cloud. For a Sume agent loop, this lets you plan at one level and wait on jobs_wait at low in the same cached conversation.

How the feature works

The effort page describes the new level taking effect from the next user turn and holding until a later message changes it. Everything before that message is unchanged, so the cached prefix still matches. Without the beta value, the request returns a 400 with Extra inputs are not permitted.

{
  "model": "claude-haiku-5-5",
  "max_tokens": 4096,
  "output_config": {"effort": "medium"},
  "messages": [
    {"role": "user", "content": "Submit the render, then wait."},
    {"role": "assistant", "content": "Submitted. Waiting."},
    {"role": "system", "content": [], "output_config": {"effort": "low"}},
    {"role": "user", "content": "Poll the job once more."}
  ]
}

Why it matters for a long job

Changing the top-level output_config.effort between requests restarts the cache, according to the same page, which advises holding it constant in cached conversations. A Sume loop that submits a render and then polls for minutes is a cached conversation: the tool list and earlier turns are a stable prefix, and a poll step costs little. A per-message change is how you vary effort there without paying to rebuild the prefix.

Per-message effort and Haiku 5.5, Claude API docs, read 2026-10-08
ItemFact
Beta headermid-conversation-output-config-2026-07-01
Haiku 5.5 on Claude API and Google CloudSupported
Haiku 5.5 with thinking disabledA differing per-message effort returns 400
Top-level effort changed between requestsDoes not keep earlier cached prefixes
Cache hit price, Haiku 5.5 up to 100k prompt$0.01 per million tokens

The Haiku 5.5 limit

If you set thinking: {"type": "disabled"} on Haiku 5.5, a per-message effort that differs from the level in effect returns a 400. To vary effort per turn, use adaptive thinking, which means omitting the thinking field or sending adaptive. Do not pass adaptive as an effort value; it is a thinking mode.

Sume's side does not care which effort produced the call. The same idempotency_key rule applies on a retry, which matters because a model at low that repeats a submit with a new key creates a second paid job. Reuse the key when you retry the same intent, and poll with jobs_wait as described in jobs_wait for long video jobs.

When not to bother

If a run is short, a single top-level effort is simpler, and a cache that expires after five minutes gives you little to protect. The feature earns its place in sessions that run for many turns with a large stable prefix, which is what a Sume agent with a full tool list and a long render looks like.

Remember it is a beta. Anthropic's page lists the supported models and platforms, and says Amazon Bedrock supports a narrower set than the Claude API and Google Cloud. Check the page for your platform before you depend on it, and keep a fallback that holds effort constant.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume