Sonnet 5.5 thinking disabled is a 400 at effort high: Sume
On Sonnet 5.5, thinking disabled is a 400 at effort high or below; use between_tools. Sume's Claude rows expose no thinking or effort fields.

On Claude Sonnet 5.5, switching off up-front thinking at effort high or below now means thinking: {"type": "between_tools"}; the old "disabled" value is rejected. If your Messages request still sends disabled, expect a 400 after you change the model id to claude-sonnet-5-5. Through Sume none of this reaches you, because Sume's Claude rows expose no thinking or effort controls to set.
The behaviour is listed under breaking changes from Claude Sonnet 5 in the Claude Platform release notes, entry of September 28, 2026, read 2026-10-03. Sume's side comes from the model catalog source and Calling a Format.
What are the Sonnet 5.5 breaking changes?
The notes list five changes against Sonnet 5, collected here with the fix each implies.
| Change | What breaks | Fix |
|---|---|---|
| Up-front thinking off | thinking disabled is no longer accepted at effort high or below | Send thinking type between_tools |
| Forced tool use | tool_choice any or tool returns 400 | Use auto with strict tool use |
| Thinking blocks | Tied to the model and conversation | Do not replay them to another model |
| Computer use tool | Earlier computer_20251124 tool not accepted on Claude API and Google Cloud | Move to the newer toolset |
| Advisor tool | Opus 4.8, 4.7 and Sonnet 5 rejected as advisors | Name a supported advisor |
Which of these can an agent author hit on Sume?
None of the request-shape errors, and that is by design of the catalog. The comment above the OpenRouter parameter list in the model catalog says those rows ship no parameters in v1: the hop is proxy-only and the per-vendor reasoning contract behind it is unverified, so an Effort menu would be a guess. Sonnet 5.5 and Opus 5.5 are both OpenRouter-routed rows, so the picker shows no effort or thinking option for them.
The cost of that simplicity is control. You cannot ask Sonnet 5.5 for lower effort on a cheap classification step through Sume. If a step needs a dial, GPT-6.1 Sol and Grok 4.7 rows do carry Effort and Thinking parameters in the catalog, though note that the GPT row's Fast axis does not apply to Grok.
What does the forced-tool change mean for video agents?
Many agent loops force a tool call to guarantee a structured step: for example, always call a render tool next. With tool_choice of any or tool returning 400 on Sonnet 5.5, that pattern needs to become auto plus strict tool use, and a prompt that says which tool is expected. A direct integration should also handle the case where the model answers in text instead of calling the tool, which forced use used to rule out.
On Sume, the orchestrator's tool choices are inside the agent, so the same concern does not appear in your request body. What you do control is the cap on spend and the model id.
curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sonnet-55-001" \
-d '{"instruction":"Three hook ideas for a vitamin C serum",
"model":"openrouter/anthropic-claude-sonnet-5-5"}'How do I confirm which model answered?
Read the receipt. The Format runs page documents a model field on the run receipt: the catalog id the orchestrator ran on. Log it with the run id so a later billing question or quality regression can be tied to a model, not to a guess.
What should a direct integration test first?
Before moving traffic to claude-sonnet-5-5, replay a small set of real requests against it and read the status codes, not the text. Any 400 is a request-shape problem from the table above and is cheap to fix. Only after the codes are clean is it worth comparing answer quality.
Keep the old and new request builders side by side behind a flag so you can roll back by flipping the model id and the thinking value together. Changing one without the other is the usual cause of a confusing partial failure.
Sources
Related posts
More in Developers
- How to compare AI video models fairly: one prompt, three models, 480p
Submit one prompt to Seedance 2.0 Mini, Wan 3.0 and MiniMax H3 at 480p and 5 seconds on Sume, then judge the clips blind. A runnable Python test harness.
- Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.
- DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.
- A dry-run flag for Sume API calls: print the request, skip the spend
Add DRY_RUN to code that calls the Sume API: build the body, key and spend cap, print them, and send nothing. Review a batch before it costs money.
Written by Sume