Haiku 5.5 returns 400 for thinking disabled at xhigh: a Sume agent fix

Claude Haiku 5.5 rejects thinking disabled at xhigh or max effort. How that 400 shows up in an agent that calls Sume, and how to set the pair.

5 min readSume
All posts

If a request to Claude Haiku 5.5 sets thinking: {"type": "disabled"} together with effort of xhigh or max, the API returns a 400. The fix is to omit the thinking field or send {"type": "adaptive"} at those levels, or to lower effort to high or below if you want thinking off. This matters for a Sume agent wrapper because a generic client often adds a fixed thinking setting to every call, and a tool loop that raises effort for one hard step can then fail before any Sume tool runs.

The rules on the vendor page

Anthropic's effort page states that Haiku 5.5 has thinking on by default, that thinking counts toward max_tokens, and that thinking: disabled is allowed at high effort or below. At xhigh or max it returns a 400, so adaptive thinking is the way to use those levels.

Haiku 5.5 thinking and effort combinations, Claude API docs, read 2026-10-08
thinking settingeffort low to higheffort xhigh or max
Omitted (default)Thinking onThinking on
adaptiveAllowedAllowed
disabledAllowed400 error

How the failure looks in a Sume loop

The 400 comes from the model API, so no Sume call was made and nothing was billed by Sume. A loop that treats any error as retryable may re-send the same request and hit the same 400 forever. Separate model errors from Sume tool errors: a Sume tool error such as insufficient_scope or an admission refusal belongs to the tool result, while a model 400 is a client bug.

One related fact from Sume's own catalog: Haiku 5.5 is the replacement for Haiku 4.5 in Sume's in-app agent model list, and the retired 4.5 row maps onto it. That is a statement about Sume's chat model picker, not about the model field of its API.

A safe default

For the dotted and dashed model-id naming problem that sits next to this one, see the Haiku 4.5 dashed alias 400.

  • Pick effort per step, and derive thinking from it: omit it above high, and send disabled only at high or below.
  • Never send both a fixed disabled and a variable effort.
  • Leave max_tokens large enough for thinking, since it counts toward the limit.
  • Log the HTTP status of the model call apart from the MCP tool result.

A small helper that cannot produce the bad pair

The simplest fix is to compute the two settings together in one function, so no code path can set a fixed disabled beside a high effort. Anthropic's rule is that disabled is accepted at high or below and rejected at xhigh and max, so the helper encodes that and nothing else.

Keep the helper in your client, not in a prompt. A model cannot be asked to avoid a 400 it never sees, and the error arrives before any tool runs. Write one test per level so a later change to the defaults shows up in CI, not in a paid run.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume