SLATEMOTH / ANALYSIS
A model upgrade changes the integration contract, not just the model ID
Claude Sonnet 5.5 shows why agents need renewed checks for thinking, tools, conversation continuity and refusals before a model switch ships.
The model name is only one line of the change
Anthropic released Claude Sonnet 5.5 on September 28, reporting better speed and task efficiency. For teams that construct Messages API requests themselves, the migration guide deserves as much attention as the performance claims. It describes changes to accepted parameters and conversation handling. Claude Managed Agents users, by contrast, are told they only need to update the model name. The analysis below concerns custom API integrations. [1] [2]
A single-turn smoke test can miss two different problems: a request that fails immediately, and a request that succeeds with different context. These need separate acceptance criteria. A successful HTTP response establishes that the service answered; it does not establish that the application preserved its workflow. That distinction should shape the migration plan.
Turning off up-front thinking has a new shape
Sonnet 5 accepted thinking: {"type": "disabled"}. Sonnet 5.5 rejects that setting with a 400 error. Its lowest thinking setting is between_tools, available at low, medium and high effort; xhigh and max reject it. This changes a request contract, not merely a tuning preference. Teams that used disabled thinking to control latency should choose the new setting deliberately and remeasure the workflow rather than assuming the old request will carry over. [2]
Between-tools mode removes up-front extended thinking, but progress updates between tool calls may still appear as thinking blocks. A tool loop must return those blocks intact with the assistant message. A reader that assumes the first content block is text, or reconstructs an assistant turn from text and tool calls alone, needs review. The safe unit to parse is a content block's type; the safe unit to replay is the assistant turn as returned. [2] [6]
A valid tool argument is not a guaranteed tool call
Sonnet 5.5 rejects the forced tool_choice values tool and any. Anthropic advises auto with strict tool schemas where supported. Strictness can constrain the shape of a tool call that occurs; auto still lets the model answer without making one. Amazon Bedrock does not offer strict tool use for this model, so applications there must validate tool input themselves. [2] [4]
That distinction is operational. If a workflow must fetch the current record before answering, a well-formed answer is insufficient evidence that the fetch happened. Mark the step complete only after the required tool result is present and validated. Put that rule in the application, then test both paths: a valid tool call and a direct answer when the call was required. This is an architectural inference from the documented tool behavior, not a claim about how often Sonnet 5.5 skips tools.
HTTP 200 can still conceal lost continuity
Sonnet 5.5 thinking blocks can be reused only by the account that produced them or a linked account. If an unrelated account sends one back, the API drops it before inference and the request succeeds; without the relevant diagnostic header, the drop is silent. Model switches can also make a block unreadable. A router's successful response therefore does not prove the model received the earlier reasoning. [3]
There is a separate prefix rule: earlier system instructions, tools and messages must remain unchanged when a signed thinking block is replayed. Accounts created on or after August 31, 2026 have this check enforced by default; older accounts have different default enforcement, so an error-free test on one key is not universal proof. Preserve the conversation as an append-only history, and exercise save-and-resume, tool changes, client-side trimming and route changes during migration. Where supported, inspect input_transformations to see whether blocks were dropped. [3]
A refusal is an outcome to handle, not a timeout to retry
The migration guide also calls for handling refusal outcomes. A refusal belongs in the product's response and review logic; blindly replaying it as though a network request failed confuses a policy result with an outage. Anthropic documents a scoped, optional server-side fallback for certain refusal categories on the Claude API, but that does not make every refusal retryable or promise the same behavior on every host. [5]
For acceptance testing, record the actual stop reason, whether an officially configured fallback occurred, and what the user saw. This keeps safety handling observable without treating a lower-level retry as a way around a refusal. It also prevents a dashboard of successful HTTP responses from hiding a change in the answers users receive.
Test the workflow that will run in production
We propose five acceptance paths: an old disabled-thinking request; a formerly forced tool call; a multi-turn tool exchange that replays thinking blocks; a saved conversation resumed through the real account and routing setup; and a refusal handled by the product. Check HTTP status, content-block types, tool execution and validation, continuity diagnostics where available, and the final user-visible state. Include a long session and a route change, since a single-turn smoke test cannot expose history errors.
Only then compare accepted-task rate, elapsed time and cost per accepted task at a stated effort level. Anthropic's speed and cost claims are useful hypotheses for that experiment, not substitutes for it. The broader lesson extends beyond this release: a model is part of a protocol between the application, tools, history and users. Upgrading it is complete when that protocol still delivers the intended result. [1]
Sources and verification
- Anthropic · Introducing Claude Sonnet 5.5, September 28, 2026
- Claude Platform Docs · Migrating to Claude Sonnet 5.5; accessed September 29, 2026
- Claude Platform Docs · Preserved thinking; accessed September 29, 2026
- Claude Platform Docs · Strict tool use; accessed September 29, 2026
- Claude Platform Docs · Refusals and fallback; accessed September 29, 2026
- Claude Platform Docs · Thinking in tool and multi-turn workflows; accessed September 29, 2026