Claude thinking blocks are model-bound: one model per thread

Sonnet 5.5 and Fable 5.1 thinking blocks only work for the model (and account) that made them. Why an agent thread should keep one model, and how to log it.

5 min readSume
All posts

If you switch Claude models in the middle of an agent thread, the earlier thinking blocks may stop working: Anthropic's September release notes say Sonnet 5.5 thinking blocks are tied to the model and the conversation, and only work in the account that produced them or linked accounts. The safe habit for a long video-agent thread is one model from first turn to last, and a log line that says which.

The statements below come from the Claude Platform release notes, entries of September 1, 14 and 28, 2026, read 2026-10-03. They describe the Claude API; this post says nothing about how any other host replays thinking.

What do the release notes say?

Three entries bear on it.

Thinking-block rules in the Claude release notes (read 2026-10-03)
DateModelRule
Sep 1Fable 5.1 and Mythos 5.1Thinking blocks preserved only for the model that produced them or newer; prefix mismatch checking is default for accounts created on or after Aug 31
Sep 14Beta, thinking-binding-controls-2026-08-01input_transformations on the response lists thinking_mismatch_allowed entries naming blocks that failed the prefix check
Sep 28Sonnet 5.5Thinking blocks tied to model and conversation; only work in the account that produced them or linked accounts

Why does this matter for a video agent?

Agent threads are long and they get handed around. A thread that starts on a cheap model for the script and moves to a stronger one for the shot review carries earlier thinking along with it. If those blocks are bound to the first model, the switch can fail or force the history to be rebuilt without them. Either outcome costs a turn.

Two practical rules follow. Pin the model id when you create the thread and carry it on every continuation, and if you want a different model for a different phase, start a new thread and pass the earlier result as plain text instead of replaying the history. Passing a summary also keeps the new thread short, which is cheaper on any model.

How does Sume handle the model on a continued run?

The Formats API makes the model an explicit field of each request. Calling a Format says model is the Agents catalog id for the LLM that orchestrates the run, and that the receipt echoes the id that ran; Format runs shows previous_run_id continuing an earlier run as another turn of the same conversation. Nothing in those pages says a continuation inherits the earlier model, so send model on every turn rather than assuming it.

The catalog also remaps retired Claude picks: Sonnet 5 and Opus 5 are marked retired onto Sonnet 5.5 and Opus 5.5, kept only so persisted rows keep their exact id. I did not find a statement of what Sume does with thinking blocks across a model change, so I do not claim one. The safe reading is: keep the id constant and log the receipt.

curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: turn-2-001" \
  -d '{"instruction":"Tighten hook 2",
       "previous_run_id":"'"$PREV_RUN_ID"'",
       "model":"openrouter/anthropic-claude-sonnet-5-5"}'

What should I log?

Keep four fields per turn: the run id, the model echoed on the receipt, the previous_run_id, and usage.debited_usd_micros. With those, a thread that suddenly behaves differently can be checked against the one thing most likely to have changed, the model, before anyone blames the prompt.

What if I really need two models in one job?

Split the job at a natural seam and pass text across it. A shot list produced by one model is just text; the second model reads it as input and starts its own thread with its own thinking. Nothing bound to the first model needs to survive the handoff.

That pattern also gives you a clean audit trail: two receipts, two model ids, two costs. When a result looks wrong you can tell which stage produced it, which is much harder when a single thread has quietly changed models partway through.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume