What Claude Opus 5.5 and Claude Sonnet 5.5 reject, what they now bill, and the exact code changes that get a Claude 4.x or 5 integration back to green
Claude Opus 5.5 costs 20% less per token than Claude Opus 5, and four request shapes that Opus 5 accepted on September 21 now come back as HTTP 400.
Anthropic launched Claude Opus 5.5 on September 22, 2026 at $4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Claude Opus 5 - Claude Platform release notes. Six days later Claude Sonnet 5.5 shipped at $2 and $10, the same price as Sonnet 5, with Anthropic claiming it "runs 30%+ faster, and costs up to 30% less for most work" - Anthropic's Sonnet 5.5 announcement. On the benchmarks that matter for production agents, both models clear their predecessors by wide margins. Neither one is a drop-in replacement.
Both models reject forced tool use (tool_choice of any or tool), both reject thinking: {"type": "disabled"}, both bind thinking blocks to the model and the conversation that produced them, and both refuse the older computer_20251124 tool on the Claude API and Google Cloud. Sonnet 5.5 adds a fifth break around the advisor tool. And two days after the Opus launch, a change that broke nothing but costs money landed: Anthropic resumed billing refusals that arrive before any output when the refusal category is bio, frontier_llm, or reasoning_extraction - How refusals are billed.
The problem is where these knobs lived. Forcing a tool call was the most common way to get structured JSON out of Claude for two years, and plenty of libraries still implement "structured output" exactly that way. Turning thinking off was how latency-sensitive routes stayed fast. Rewriting old turns was how harnesses kept their context small. Code that leaned on any of these fails on its first request to the new models, or, in the case of history edits on older accounts, succeeds while quietly losing the reasoning it paid for.
This guide covers each breaking change with the exact error text, before-and-after code in Python and TypeScript, and the edge cases the docs bury (Amazon Bedrock has no strict tool use for either 5.5 model, for example). It then walks through the changes that fail no request, what the ecosystem broke, how to choose between Opus 5.5 and Sonnet 5.5, the real cost math, and a runbook. If you searched for a Claude Opus 5.5 migration checklist, it is section 10, but read the first five sections before running it: the checklist only makes sense once you know which of your requests carry each problem.
Contents
- What Shipped on September 22 and 28, and Why the Knobs Disappeared
- Breaking Change 1: Thinking Can't Be Turned Off
- Breaking Change 2: Forced tool_choice Returns HTTP 400
- Breaking Change 3: Thinking Blocks Are Bound to the Model and the Conversation
- Breaking Change 4: Computer Use Needs the New Toolset (and Sonnet's Fifth Break)
- The Changes That Fail No Request: Effort, Progress Notes and Refusal Billing
- What the Ecosystem Broke: Frameworks, Gateways and Claude Code
- Opus 5.5 or Sonnet 5.5: Where to Land
- The Cost Math of Migrating
- A Claude Opus 5.5 Migration Runbook You Can Ship This Week
- What These Changes Signal About Building on Frontier Models
- Conclusion: A Decision Framework
Where to Land: Six Options Scored
Before the mechanics, the decision most teams are actually making: migrate now, and to which model, or stay put for a cycle. The table scores six realistic landing spots on four criteria. Capability (30%) uses independent measurements where they exist (the official Terminal-Bench 4.0 leaderboard and the Artificial Analysis Intelligence Index) and Anthropic's launch tables where they do not. Cost per task (30%) combines list price, cache-read price and Artificial Analysis's measured cost per task at each effort level. Migration effort (20%) counts the breaking changes you must absorb from a Claude Opus 5 or Sonnet 5 codebase. Platform fit (20%) covers the operational edges: Priority Tier, zero data retention, Amazon Bedrock feature gaps and fast mode.
Staying on the old model scores well on migration effort by definition, which is the honest trade: it costs nothing this week and costs more per task every week after. The sort is by final score, highest first.
| # | Option | What It Is | Capability (30%) | Cost per task (30%) | Migration effort (20%) | Platform fit (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Current Opus, launched Sep 22, $4/$20 | 10 - #1 on the Terminal-Bench 4.0 leaderboard (64.85%), AA index 58 at max | 8 - $4/$20, cache reads $0.20; AA: medium matches Opus 5 at max for $1.34 vs $5.86 a task | 5 - four breaking changes from Opus 5 | 8 - all five platforms, ZDR available, fast mode; no Priority Tier | 8.0 |
| 2 | Claude Sonnet 5.5 | Current Sonnet, launched Sep 28, $2/$10 | 9 - 61.82% Terminal-Bench 4.0 leaderboard, AA index 56, 1844 GDPval-AA Elo | 8 - $2/$10, cache reads $0.20; AA: $0.59 a task at medium, $7.67 at max | 4 - five breaking changes, incl. advisor pairings | 7 - all five platforms; no strict tools on Bedrock; no Priority Tier | 7.3 |
| 3 | Stay on Claude Opus 5 | Previous Opus, still served everywhere | 7 - 51.82% Terminal-Bench 4.0 leaderboard, AA index 51 at max | 5 - $5/$25, cache reads $0.50, $5.86 a task at max (AA) | 10 - zero code changes | 7 - unchanged surface, no retirement date announced | 7.0 |
| 4 | Stay on Claude Sonnet 5 | Previous Sonnet, same $2/$10 price | 5 - 1449 GDPval-AA Elo, 34.1% CursorBench 4.0 (vendor) | 7 - same list price, fewer tokens a task than Sonnet 5.5 at max (AA) | 10 - zero code changes | 6 - no mid-conversation system messages, 1,024-token cache minimum | 6.8 |
| 5 | Stay on Claude Opus 4.8 | Older Opus that still accepts forced tools | 6 - prior generation, now a cyber re-route target | 5 - $5/$25, cache reads $0.50 | 9 - accepts forced tools and thinking off | 8 - keeps Priority Tier and fast mode, retires no sooner than May 2027 | 6.7 |
| 6 | Claude Fable 5.1 | Top widely released tier, $10/$50 | 8 - 55.8% Terminal-Bench 4.0 (vendor), AA index 53.4 | 3 - $10/$50, cache reads $0.25 | 6 - same three core breaks, no computer-tool change | 5 - 30-day retention required, no Priority Tier | 5.5 |
Sources for the cells: the official Terminal-Bench 4.0 leaderboard, Artificial Analysis's per-effort model pages for index and cost per task, the OpenRouter models API for the Fable index, Anthropic's Opus 5.5 launch page and Sonnet 5.5 page for the vendor columns, and the pricing page for prices, all read on October 6, 2026. Columns marked "vendor" are Anthropic's own numbers; the rest are independent.
What the table says, in plain terms: the two 5.5 models are close, and the gap between them depends on effort more than on the model name. Independent measurements put Opus 5.5 on top: it leads the Terminal-Bench 4.0 leaderboard and the Artificial Analysis index, and at its medium default it reaches the same index score as Opus 5 at maximum effort for less than a quarter of the cost per task. Sonnet 5.5 wins at low and medium effort, where it is the cheapest frontier option per task, and loses its edge at max effort, where Artificial Analysis measured "the highest token use we have measured". Anthropic itself says "for the hardest long-horizon work, an Opus model is the better choice" - Prompting Claude Sonnet 5.5. Staying put is a legitimate short-term call for a team mid-launch (Opus 4.8, covered in our Opus 4.8 guide, still accepts every request shape the 5.5 models reject), but the scores show it is a deferral, not a strategy. Section 8 goes deeper on the choice; the next seven sections are about making either migration work.
1. What Shipped on September 22 and 28, and Why the Knobs Disappeared
The 5.5 family did not arrive out of nowhere. Every Opus 5.5 breaking change except the computer toolset rule already existed on Anthropic's top tier: always-on thinking since Claude Fable 5, and forced-tool rejection plus "preserved thinking" since Claude Fable 5.1 launched on September 1 - release notes. Our Claude Fable 5 guide covered the first of those shifts when it happened. Teams that run Fable 5.1 already did most of this migration. Teams on Opus 5 or Sonnet 5 skipped it, because those changes had not yet reached the Opus and Sonnet lines. On September 22 and 28 they did.
Anthropic's launch framed Opus 5.5 as "the first model in our new Claude 5.5 family", one that "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5" - Anthropic's Opus 5.5 launch page. The same page notes, in a single line, that the model "is also no longer available with 'thinking' mode switched off", which is the first breaking change in this guide.
The timing matters because the changes are cumulative and account-dated. The history-editing check that defines preserved thinking is enforced by default only for API accounts created on or after August 31, 2026, 00:00 UTC. An older account can run a harness that edits history for months without an error, while a brand-new account (a new customer's workspace, a new staging org, a fresh CI account) gets HTTP 400 on the same code - Preserved thinking. That asymmetry is the single most likely source of "works on my machine" bugs in this migration.
| Date (2026) | Change | Who it affects |
|---|---|---|
| Aug 31 | Accounts created from this date get the preserved-thinking check by default | New orgs, workspaces, CI accounts |
| Sep 1 | Claude Fable 5.1 rejects forced tool_choice, binds thinking blocks | Fable users |
| Sep 22 | Claude Opus 5.5 launches: thinking always on, four breaking changes | Opus 5 / 4.x users |
| Sep 24 | Pre-output refusals billed again for bio, frontier_llm, reasoning_extraction | Everyone on classifier models |
| Sep 28 | Claude Sonnet 5.5 launches: five breaking changes, between_tools | Sonnet 5 / 4.x and Haiku users |
| Sep 30 | Claude Sonnet 4.5 deprecated, retiring on the Claude API on Nov 30 | Sonnet 4.5 users |
Prices moved the other way from the API surface. The Opus line got cheaper, the Sonnet line held, and cache reads collapsed: Opus 5.5 bills cache hits at 0.05x the base input price ($0.20 per million tokens), where every other model except Fable 5.1 bills 0.1x - pricing. The chart shows list prices for the models a migrating team is choosing between.
Sonnet 5's $2/$10 price was announced as introductory pricing through August 31, with a scheduled rise to $3/$15; Anthropic dropped that increase, and $2/$10 is now the standard Sonnet 5 price as well - model pricing. So Sonnet 5.5 does not save money per token over Sonnet 5. It saves money per task, which is a claim you have to measure on your own traffic rather than take from a launch post.
Why Anthropic took the knobs away
Anthropic does not publish one paragraph explaining why forced tool use and the thinking switch disappeared, so this is a reading of the structure, not a quote. Start with what the knobs did. Forced tool_choice let the API, not the model, decide the first action of a turn. Disabled thinking let the caller decide that the model would not reason before acting. Assistant prefill, removed back on the 4.6 generation, let the caller write the first words of the answer. All three are the same kind of control: the request overrides the model's own judgment about how to start.
The 5.5 models are built around the opposite assumption. Thinking is always on, the model decides how much to think, and every action it takes is preceded by reasoning the API can verify. The pricing page makes this concrete in an odd place: its table of tool-use system prompt tokens lists 286 tokens for auto and none on both Opus 5.5 and Sonnet 5.5, and leaves the any/tool column empty, where Opus 5 lists 286 and 406 - tool use pricing. As far as Anthropic's own pricing table shows, the forced path is not deprecated on these models; it does not exist in how they are prompted.
The control is moving from the request to the output contract. What you can still constrain is the shape of what comes back (strict tool schemas, structured outputs) and how much the model spends getting there (effort, max_tokens, task budgets). What you can no longer do is reach into the middle of the model's decision. For a builder, that suggests a mental model shift: treat the model less like a function with flags and more like a contractor with a brief, a deliverable format and a budget.
The second structural thread is protecting the reasoning itself. Anthropic's docs say plainly that preserved thinking "guards against distillation" - Preserved thinking. Read alongside the other changes, a pattern appears: thinking blocks are signed and valid only in the conversation (and, for Sonnet 5.5, the account) that produced them; prompts that ask the model to write its reasoning into the response get declined as reasoning_extraction; a frontier_llm category declines help building competing models; and pre-output refusals in exactly those categories are billed "to disrupt attempts to circumvent Anthropic's safeguards at scale" - refusal billing. The reasoning is the asset, and the API surface is being redesigned to keep it from leaving.
Why this matters for you: neither thread is likely to reverse. Every model Anthropic shipped in September carries the same constraints, so the right response is not a patch for one model but an integration that assumes them. How to apply it: isolate model-specific request shapes behind one adapter, make your conversation history append-only, and replace every forced call with an output contract. The next four sections do exactly that, one breaking change at a time.
2. Breaking Change 1: Thinking Can't Be Turned Off
On Claude Opus 5, thinking was on by default but could be disabled at high effort or below. On Claude Opus 5.5, thinking is always on. Both thinking: {"type": "disabled"} and the old manual budget form ({"type": "enabled", "budget_tokens": N}) return a 400 invalid_request_error at every effort level, with the message "thinking.type.disabled" is not supported for this model. (or "thinking.type.enabled" ... for the budget form) - Opus 5.5 migration guide. Omitting the field and sending {"type": "adaptive"} are equivalent, and effort becomes the only dial.
Claude Sonnet 5.5 is slightly different, and the difference matters for latency-sensitive teams. It also rejects disabled, but it offers a new lowest setting, thinking: {"type": "between_tools"}, under which the model does no up-front thinking and only writes short progress notes between tool calls. Without tools, the response is plain text, exactly as disabled was on Sonnet 5. The 400 message for disabled even tells you the fix: To turn thinking off on this model, send "thinking": {"type": "between_tools"} instead of {"type": "disabled"}. - Sonnet 5.5 migration guide.
The before and after
The Opus change is a deletion plus a decision. Delete the thinking field, then choose an effort level. Where you disabled thinking to save tokens or time, start at low, measure latency and quality on real traffic, and move to medium if quality drops. Anthropic's own prompting guide suggests a system prompt line, "Answer directly without deliberating.", when time to first token still matters after that, with the warning that less thinking can cost accuracy - Prompting Claude Opus 5.5.
import anthropic
client = anthropic.Anthropic()
# Before: accepted on Claude Opus 5, HTTP 400 on Claude Opus 5.5
client.messages.create(
model="claude-opus-5",
max_tokens=1024,
thinking={"type": "disabled"},
messages= [{"role": "user", "content": "Classify this support ticket: ..."}],
)
# After: thinking is always on, effort is the dial
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=8000, # room for thinking plus the reply
output_config={"effort": "low"}, # start low where thinking was off
messages= [{"role": "user", "content": "Classify this support ticket: ..."}],
)
reply = "".join(block.text for block in response.content if block.type == "text")
On Sonnet 5.5 you have two migration paths, and the docs recommend trying the first one before the second. Adaptive thinking at low effort keeps thinking short and skips it on most simple requests. If a route genuinely must stay thinking-off, between_tools is the replacement, with three restrictions that will bite anyone who copies a Sonnet 5 config verbatim.
- Effort
highor below. Atxhighormax,between_toolsreturns a 400. - No other fields.
display,budget_tokensorblock_bindingalongside it is a 400. - No per-turn effort changes. A per-message effort change is rejected under
between_tools. - Sonnet 5.5 only. Every other model rejects the value, so a retry on another model must drop it.
That last restriction is the one that breaks routers and retry wrappers. A client-side retry that re-sends the same body to Claude Sonnet 5 or Opus 5.5 after a refusal will fail with "thinking.type.between_tools" is not supported for this model. Server-side fallback under the server-side-fallback-2026-07-01 header handles this for you (a between_tools request that falls back to Sonnet 5 runs there with thinking disabled), but any retry code you wrote yourself needs an explicit model-aware rewrite - Refusals and fallback.
# Sonnet 5.5, preferred: adaptive thinking (the default) at low effort
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=8000,
output_config={"effort": "low"},
messages= [{"role": "user", "content": "..."}],
)
# Sonnet 5.5, when a route must stay thinking-off
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=8000,
thinking={"type": "between_tools"}, # high effort or below, no other fields
output_config={"effort": "high"},
messages= [{"role": "user", "content": "..."}],
)
The response-shape bugs that follow
Turning thinking on changes the shape of every response, and that is where most silent breakage hides. A response can now begin with one or more thinking blocks before the first text block, and under the default display: "omitted" those blocks carry an empty thinking string plus a signature. Code that reads response.content [0].text, or a stream handler that assumes the first content_block_start is text, breaks on the first response that thinks. The fix is boring and universal: select blocks by type, never by position.
The second bug is budget. Thinking tokens count toward max_tokens and are billed as output tokens even when their text is not returned to you - Thinking. A route that ran with thinking off at max_tokens: 256 can now cut replies off mid-sentence, because the model spent the budget thinking. Anthropic's guidance for long agentic turns is to start at 64K and tune, and its Sonnet 5.5 guidance for agentic coding is 128,000 (the model's maximum) with streaming, since SDKs require streaming above roughly 21,333 tokens.
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 8000,
output_config: { effort: "low" },
messages: [{ role: "user", content: "Summarize this incident report: ..." }],
});
// Read by block type: thinking blocks can come first, with empty text
const reply = response.content
.filter((block): block is Anthropic.TextBlock => block.type === "text")
.map((block) => block.text)
.join("");
if (response.stop_reason === "max_tokens") {
// Thinking can consume the budget on 5.5 models: treat as failed and retry larger
}
There is also a prompt-level cleanup. If your Opus 5 integration ran with thinking off, it may contain instructions that stood in for thinking ("think step by step in a <reasoning> section before answering"). On the 5.5 models, those instructions are not just redundant: a prompt that pushes the model to reproduce its reasoning in the response text can be declined with the reasoning_extraction refusal category. Remove them, set display: "summarized" if you need to see the reasoning, and read it from the thinking blocks. Also delete any "don't think" rule; the model cannot comply, and such rules make internal XML tags more likely to leak into output.
Why this matters: this break is the one most likely to slip into production unnoticed, because the obvious 400 is easy to fix and the follow-on shape and budget bugs are not. How to apply it: grep for every disabled, every budget_tokens, every content [0], every small max_tokens, and every prompt that asks for written-out reasoning, and fix them in the same change. Then measure time to first token on real traffic at low effort before deciding whether a route needs between_tools on Sonnet 5.5 or a different model entirely.
3. Breaking Change 2: Forced tool_choice Returns HTTP 400
This is the break with the widest blast radius. On Opus 5.5 and Sonnet 5.5, tool_choice: {"type": "any"} and tool_choice: {"type": "tool", "name": "..."} return a 400 with tool_choice: type "tool" and "any" are not supported for this model. The rejection applies to the token-counting endpoint as well as the Messages API, so even a pre-flight count_tokens call with a forced tool fails - Opus 5.5 migration guide. {"type": "auto"} (the default) and {"type": "none"} are unchanged.
The reason this hurts is historical. Before structured outputs existed, the reliable way to get JSON out of Claude was to define a tool whose schema was your desired output and force the model to call it. That pattern is baked into a generation of code: extraction pipelines, classifiers, routers that force a route_to tool, evaluation harnesses that force a grade tool, and many framework wrappers whose with_structured_output() or generateObject() helpers do the forcing for you. Section 7 covers the frameworks; this section covers the code you own.
Migrate by intent, not by syntax
There is no single replacement for a forced call, because forced calls were used for two different jobs. The first job is extraction: you never wanted a tool call at all, you wanted a typed object. The second job is action: the tool call is the thing you need to happen, such as fetching the weather before answering or writing a record. Each job has a different replacement, and mixing them up produces code that works in tests and fails at the edges.
The distinction is worth making explicitly in code review, because the two jobs look identical in a request body. Both send a tool definition and a forced tool_choice; only the code that reads the response reveals which job it was doing. A useful test: if your code discards the tool_use block after reading its input and never returns a tool_result, it was extraction wearing a tool costume. If it executes something and feeds the result back, it was an action. That one question sorts most call sites in a large codebase in an afternoon.
- Extraction moves to structured outputs (
output_config.format), which constrains the response text itself. - Action moves to
tool_choice: autoplusstrict: true, with the tool named in the prompt. - Verification becomes your job:
autodoes not guarantee a call, so check and retry. - Bedrock is the exception: strict tools and structured outputs are not available there for these models.
The verification point is the one teams skip. With a forced call, the API guaranteed a tool_use block. With auto, the model can answer in text instead, and on a weak prompt it sometimes will. Strict mode guarantees that if the tool is called, its input matches the schema (grammar-constrained sampling, per the strict tool use docs), but it says nothing about whether the call happens. So the replacement for a forced call is always a triplet: steer in the prompt, constrain the schema, verify the result.
Action calls: auto, strict and a retry
For action calls, the docs' own before-and-after is the minimal version: mark the tool strict, send auto, and append "Use the get_weather tool" to the user message. Production code needs the check. The helper below asks for a specific tool, verifies that the model called it, and if it did not, appends the model's turn plus a one-line nudge and tries once more. The retry appends rather than edits, which keeps the history valid for preserved thinking (section 4).
import anthropic
client = anthropic.Anthropic()
class NoToolCall(Exception):
pass
def call_tool(tools: list [dict], messages: list [dict], tool_name: str, attempts: int = 2):
strict_tools = [{**tool, "strict": True} for tool in tools] # schemas need additionalProperties: false
history = list(messages)
for _ in range(attempts):
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=8000,
tools=strict_tools,
tool_choice={"type": "auto"},
messages=history,
)
if response.stop_reason == "refusal":
raise RuntimeError(f"refused: {response.stop_details}")
calls = [b for b in response.content if b.type == "tool_use" and b.name == tool_name]
if calls:
return calls [0], response
# Append, never edit: earlier turns and their thinking blocks stay byte-identical
history += [
{"role": "assistant", "content": response.content},
{"role": "user", "content": f"Call the {tool_name} tool now, using the details above."},
]
raise NoToolCall(tool_name)
Two details in that helper are not obvious. First, the tool definitions themselves must be declared from the first request of a session and never edited, because changing the tools array mid-conversation invalidates later thinking blocks; if you need to add or remove tools mid-session, the API now has tool_addition and tool_removal blocks for that. Second, strict mode has limits: at most 20 strict tools per request, at most 24 optional parameters across all strict schemas, at most 16 union-typed parameters, every object must set additionalProperties: false, and keywords like minimum, maximum, minLength and recursive schemas are not supported - JSON Schema limitations. The first request with a new schema also pays a grammar-compilation latency hit; compiled grammars are cached for 24 hours from last use.
If you had disable_parallel_tool_use: true alongside a forced call to get exactly one call, note the semantic shift: with auto, that flag can cap the turn at one call but cannot guarantee that one happens. The verification step is what turns "at most one" back into "exactly one".
Extraction calls: structured outputs
For extraction, structured outputs are the cleaner replacement, because the JSON arrives as the response text with no tool round trip and no fake tool definition. Set output_config.format to a JSON schema, and the response text is guaranteed to match it. On Claude Sonnet 5.5 there is a subtlety worth knowing: on tasks that need a few steps of working out (totaling figures, applying a rule, ranking items), the model often answers without thinking first at low and medium effort, and with structured outputs it has nowhere else to reason. Anthropic's fix is a single system prompt line, "Think the problem through before you answer.", which at high effort brings accuracy close to xhigh for a modest token increase - Prompting Claude Sonnet 5.5.
import json
schema = {
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["billing", "bug", "how_to", "other"]},
"urgent": {"type": "boolean"},
"summary": {"type": "string"},
},
"required": ["category", "urgent", "summary"],
"additionalProperties": False,
}
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
system="Think the problem through before you answer.",
output_config={"effort": "high", "format": {"type": "json_schema", "schema": schema}},
messages= [{"role": "user", "content": ticket_text}],
)
if response.stop_reason == "max_tokens":
raise RuntimeError("thinking ran out the budget; retry with a larger max_tokens")
ticket = json.loads(next(b.text for b in response.content if b.type == "text"))
The max_tokens check in that snippet is not decoration. Anthropic notes that with structured outputs at low and medium effort, Sonnet 5.5 "occasionally keeps thinking until it reaches max_tokens", and a response that stopped there should be treated as failed even if its text parses. If you cannot use structured outputs at all, the documented fallback is to ask for JSON in the prompt and parse the last complete JSON value in the response, because the model often works the problem out in prose first and writes the JSON at the end.
There is one case where structured outputs are the wrong replacement, and it comes from field data rather than the docs. Inside an agent loop that also has tools, a request that combines native structured outputs with a tool list can leave the model calling tools instead of producing the final object. A Vercel AI SDK user replayed the same agent task 20 times on Sonnet 5.5 with native structured output and got a valid final answer 0 times out of 20, while Opus 5.5 managed 20 of 20; switching to tool_choice: auto with a strict final_answer tool produced 40 of 40 on both models - vercel/ai issue 21992. The rule of thumb that follows: structured outputs for single-shot extraction, a strict "final answer" tool for the last step of a tool-using agent.
The Amazon Bedrock exception
Here is the edge case most migration posts miss. On Amazon Bedrock, structured outputs, which include strict tool use, are not available for Claude Opus 5.5 or Claude Sonnet 5.5. They exist only on the legacy Bedrock integration for Opus 4.6 and earlier - structured outputs. That means a Bedrock team migrating off forced tool use loses both of its guarantees at once: the call is no longer guaranteed, and the input is no longer schema-constrained. The failure is not always loud, either. LiteLLM users found that a response_format request to Opus 5.5 on Bedrock, which the library turned into an unforced tool, came back as plain text with HTTP 200 in five of five tries - LiteLLM issue 42717. Nothing errored; the JSON simply never arrived.
The documented Bedrock pattern is auto without strict, a prompt that says when to call the tool, and schema validation in your own code. In practice that means a validation layer (Pydantic in Python, Zod in TypeScript) between the model and every tool handler, returning a tool_result with is_error: true and the validation message when input does not parse, so the model can correct itself on the next turn. Teams that run the same codebase on the Claude API and Bedrock should build that validation layer everywhere: it costs little on the Claude API, where strict mode makes it a no-op, and it is the only safety net on Bedrock.
Why this matters: forced tool use was load-bearing in a lot of code that does not look like tool code, and on Bedrock its replacement is weaker than on the Claude API. How to apply it: classify every forced call as extraction or action, move extraction to structured outputs, move action to auto plus strict plus a verify-and-retry loop, and put schema validation in front of every tool handler regardless of platform.
4. Breaking Change 3: Thinking Blocks Are Bound to the Model and the Conversation
The third break is the subtlest, because on many accounts it does not raise an error at all. Anthropic calls it preserved thinking: when a thinking or redacted_thinking block comes back in a request, the API checks its signature for two things, whether the current model is allowed to read it, and whether everything before it is byte-identical to what was sent when the block was produced - Preserved thinking. Opus 5.5 and Sonnet 5.5 both run both checks. Sonnet 5.5 adds a third: its blocks only work in the account that produced them, or an account linked to it.
The two checks fail differently, which is why this break is so easy to miss in testing. A block the model cannot read is silently dropped: the request succeeds, the dropped block is not billed, and the model simply answers without that reasoning. A block whose prefix changed either returns a 400 or is dropped, depending on your account's age and settings. Neither failure shows up in a unit test that sends one turn.
Model binding: who reads whose reasoning
Every model reads its own thinking blocks plus a fixed set of other models' blocks, and the sets are not symmetric. Opus 5.5 reads blocks from Opus 5 and earlier Opus, Sonnet and Haiku models, and on the Claude API and Google Cloud it also reads Sonnet 5.5's blocks; it does not read Fable or Mythos blocks. Sonnet 5.5 reads Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models, but not Opus 5, Opus 5.5, Fable or Mythos. And on the Claude API, Fable 5.1 and Mythos 5.1 are the only models that read Opus 5.5's blocks - Thinking: switching models.
The practical consequence lands on routers and fallbacks. A conversation that moves up keeps its reasoning: Sonnet 5.5 to Opus 5.5 on the Claude API, or Opus 5.5 to Fable 5.1 on the Claude API. A conversation that moves anywhere else loses it for that request. That includes the most common migration-era pattern, a refusal fallback from Opus 5.5 to Opus 4.8 or Opus 5, and the most common cost pattern, a router that drops easy turns from Opus 5.5 to Sonnet 5.5. The fallback model answers, the request succeeds, and nothing in the response tells you the model lost the reasoning behind the previous ten tool calls unless you opt into the thinking-binding-controls-2026-08-01 beta header, which lists each dropped block in an input_transformations array with reason: "model_binding_mismatch".
Anthropic's instruction is to keep sending the full history, thinking blocks included, and let the API drop what the current model cannot read. Do not strip blocks yourself on a model switch; the one time stripping is explicitly wrong is when you redeem a fallback credit, which requires the exact original body. If your cost router bounces conversations between tiers on every turn, the reasoning loss is a real quality cost, and section 9's cost math should include it: a cheaper model that has to re-derive context is not always cheaper per task.
Conversation binding: the history-editing check
The second check is the one that throws. A thinking block stays valid only while the top-level system prompt, the tools array and every message before the block are unchanged. Change any of them and that block, plus every later thinking block, becomes invalid. For accounts created on or after August 31, 2026, 00:00 UTC, the default behavior is a 400 invalid_request_error, usually ending with a sentence that names what changed - Preserved thinking. Older accounts are enforced only on requests that set prefix_mismatch_behavior, which means the same code can pass every test in an old org and fail on its first request in a new one. The message looks like this:
messages.1.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".
What counts as an edit is broader than most teams expect. These are the patterns that show up in real harnesses:
- Rebuilding the system prompt each turn with fresh state, a timestamp or a new instruction.
- Re-rendering context in the first user message (date, memory, project instructions).
- Shrinking old tool results or re-encoding old images in place to save tokens.
- Changing the tools array mid-session, including adding a tool the model now needs.
- Inserting a reminder for one request and deleting it on the next.
Each of those has an append-only replacement, and most of the replacements are new API features shipped in the same months as the models. The pattern behind all of them is the same: instead of rewriting the past, you add a new message that changes what happens next, so everything the model already reasoned about stays byte-identical. Anthropic's own table maps the edits to their replacements one to one. The version below keeps the beta headers, because sending the right feature with the wrong header produces its own 400, and several of these headers are dated within weeks of each other.
| Instead of | Use | Beta header |
|---|---|---|
Rebuilding the top-level system prompt | A mid-conversation role: "system" message | None |
| Re-rendering context in the first user message | Resend it unchanged, put the update in the newest turn | None |
Shrinking old tool_result content in place | Server-side clearing with clear_tool_uses_20250919 | context-management-2025-06-27 |
| Inserting a reminder, deleting it next turn | Turn-scoped system message, clear_at: "next_user_message" | mid-conversation-system-clear-at-2026-08-21 |
Editing the tools array | tool_addition and tool_removal blocks | inline-tools-2026-09-15 |
| Changing top-level effort between turns | A per-message output_config | mid-conversation-output-config-2026-07-01 |
| Summarizing old turns on the client | On-demand compaction, or simple compaction with no replayed thinking | compact-2026-09-04 |
Two of these deserve a note. The mid-conversation system message is the workhorse: it carries system-prompt authority, needs no beta header on the 5.5 models, and lets a harness change instructions without touching the cached prefix - mid-conversation system messages. And client-side compaction has a trap: "keep-tail" compaction, where you summarize old turns but keep the last few verbatim, breaks the check, because the kept assistant turns carry thinking blocks signed against the original history. Either let the API write the summary with on-demand compaction, replay no earlier thinking at all, or set drop_block so the stale blocks are dropped instead of rejected.
Audit your harness before your next new account does
The worst version of this bug ships quietly on an old account and detonates on a new one. Anthropic documents a three-step audit, and it is worth running even if your account is exempt, because it also raises your prompt-cache hit rate. Capture request bodies over a few normal turns (including one with a compaction or tool change), diff consecutive pairs, and then confirm with the API by running a real session with drop_block set and counting dropped blocks.
import anthropic
client = anthropic.Anthropic()
turns = ["List the open invoices for ACME.", "Which of those are overdue?", "Draft a reminder email."]
messages = []
for turn in turns:
messages.append({"role": "user", "content": turn})
response = client.beta.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive", "block_binding": {"prefix_mismatch_behavior": "drop_block"}},
messages=messages, # your real harness builds this; the point is to run it unchanged
betas= ["thinking-binding-controls-2026-08-01"],
)
messages.append({"role": "assistant", "content": response.content}) # echo exactly as received
dropped = response.input_transformations or []
print(f"dropped blocks: {len(dropped)}", [d.reason for d in dropped])
A clean integration prints zero dropped blocks every turn. On an older account, also watch for thinking_mismatch_allowed entries: they name blocks that failed the prefix check but were let through because your account is not enforced, which is exactly the list of edits that will become 400s on a new account. Log input_transformations on every production turn during the migration window, then decide whether to opt in to "error" explicitly so every environment behaves the same.
There is good news for one group. Anthropic states that Claude Code, claude.ai, Managed Agents and the Agent SDK already keep the prefix intact, so integrations built on those harnesses need no change for this break. It bites teams that build messages themselves: custom agent loops, LangGraph-style state machines that rewrite history, chat apps that re-render a context block, and any serializer that drops empty fields. That last one is worth checking explicitly: under the default display: "omitted", thinking blocks have an empty thinking string, and a JSON serializer that strips empty strings or unknown block types edits the prefix on every turn. So does filtering on block.type == "thinking" and silently losing redacted_thinking blocks.
Why this matters: this break turns history-management shortcuts into either 400s or silent quality loss, and which one you get depends on when your account was created. How to apply it: make every harness append-only, echo assistant turns byte-for-byte, replace each edit with its API feature, run the drop_block audit, and treat model switches as reasoning resets in your quality and cost planning.
5. Breaking Change 4: Computer Use Needs the New Toolset (and Sonnet's Fifth Break)
The fourth break only affects teams running computer use, but for them it is more than a rename. On the Claude API and Google Cloud, a tools entry of type computer_20251124 returns a 400 that begins 'claude-opus-5-5' does not support tool types: computer_20251124. and then lists the accepted types. Claude Sonnet 5.5 behaves the same way and also rejects the older computer_20250124 on every platform. On Amazon Bedrock, both 5.5 models still accept computer_20251124 with its beta header - Opus 5.5 migration guide.
The replacement is computer_toolset_20260801, which is GA on the Claude API and Google Cloud with no beta header. It is a different shape, not a different version string. The entry takes no name and no display dimensions, an optional configs map turns member tools on or off, and the model's calls arrive as separate tool_use blocks named after the member (screenshot, left_click, type, zoom and so on, 17 members in all) rather than one computer tool with an input.action field - computer use tool.
# Before: HTTP 400 on Claude Opus 5.5 and Sonnet 5.5 (Claude API, Google Cloud)
client.beta.messages.create(
model="claude-opus-5",
max_tokens=4096,
betas= ["computer-use-2025-11-24"],
tools= [{"type": "computer_20251124", "name": "computer",
"display_width_px": 1024, "display_height_px": 768}],
messages= [{"role": "user", "content": "Open the display settings."}],
)
# After: no beta header, no name, no display size
client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
tools= [{"type": "computer_toolset_20260801", "configs": {"zoom": {"enabled": True}}}],
messages= [{"role": "user", "content": "Open the display settings."}],
)
The agent loop changes more than the request. Read the action from each block's name, not input.action. Expect several tool_use blocks in one turn (a batch action), run them in order, and stop at the first failure, returning the documented halt text for the skipped ones. Return one tool_result per call in the next user message, and every result must echo "toolset_name": "computer"; a result that omits it is rejected. Only screenshot and zoom results need an image, and screenshots must already fit the model's image limits, because the toolset takes no display dimensions and the API will not downscale for you.
Three smaller gotchas round it out. The fine-grained-tool-streaming-2025-05-14 beta header returns a 400 alongside a toolset entry, so replace it with eager_input_streaming: true on the tools that need it. Declaring the toolset with its default members adds roughly 4,500 input tokens per request, and disabling zoom removes about 410 of them - computer use pricing. And on Sonnet 5.5 specifically, do not prune old screenshots on the client to save context: removing an earlier screenshot edits the prefix and invalidates every later thinking block, so use server-side tool-result clearing instead.
| Version you send today | Typical starting models | Send on Claude API and Google Cloud | Send on Amazon Bedrock |
|---|---|---|---|
computer_20251124 | Claude Sonnet 5, Sonnet 4.6, Opus 5 | computer_toolset_20260801 | computer_20251124 |
computer_20250124 | Sonnet 4.5, Haiku 4.5, Sonnet 4 | computer_toolset_20260801 | computer_20251124 |
Sonnet 5.5's fifth break: the advisor tool
Sonnet 5.5 has one breaking change Opus 5.5 does not. With the advisor tool (beta), where a cheaper executor model consults a stronger advisor model mid-task, a Sonnet 5.5 executor only accepts advisors from the current generation: Opus 5, Opus 5.5, Sonnet 5.5, Fable 5, Fable 5.1, Mythos 5 or Mythos 5.1. Pairing it with Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 or Sonnet 4.6 returns a 400 - Sonnet 5.5 migration guide.
The second half of that change is easy to miss and matters for observability. Every accepted advisor returns its advice encrypted, as an advisor_redacted_result block, so the advice text is not readable in the response. If your logging, evals or debugging relied on reading what the advisor said, that visibility is gone with a Sonnet 5.5 executor. And because the executor rejects forced tool_choice, you can no longer force a consult; you nudge it from the prompt, the same steer-and-verify pattern as section 3.
Why this matters: computer-use loops are long and expensive, and a half-migrated loop fails mid-session, not on the first request. How to apply it: migrate the toolset on Claude Opus 5 first (it accepts both forms), rewrite the loop for member-named blocks, batch actions and toolset_name, keep computer_20251124 on Bedrock, and re-pair any advisor setup before moving a Sonnet executor.
6. The Changes That Fail No Request: Effort, Progress Notes and Refusal Billing
The four breaking changes announce themselves. The changes in this section do not: every request succeeds, and the effects show up later as a quality dip, a quiet UI, or a line on the invoice. For many teams they matter more than the 400s, because nobody is looking for them.
Effort defaults moved, and the levels mean different things
On Claude Opus 5.5, the API default effort is medium; Claude Opus 5 and earlier Opus models default to high. A request that omits effort now runs one level lower than it did yesterday, and setting effort to the default is identical to omitting it - Effort. That sounds like a quality regression, but Anthropic's testing says the opposite: Opus 5.5 at medium "matches or exceeds Claude Opus 5 at high" on coding and knowledge-work evaluations, and on several coding evaluations low comes close at much lower cost. The catch runs the other way: at a given level Opus 5.5 thinks more per turn than Opus 5, especially at xhigh and max, so a team that carries over an explicit xhigh setting will see longer turns and bigger bills.
Claude Sonnet 5.5 keeps high as its API default, but the levels are recalibrated: a level does not produce the same amount of thinking it did on Sonnet 5. Anthropic's starting points are medium for agentic coding and multistep tool use (move to high for harder or longer tasks), and medium or low for chat and latency-sensitive work. One wrinkle: the Sonnet 5.5 announcement says "in Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High", so a behavior you observed in Claude Code is not what an API call with no effort gets.
| Model | API default effort | Anthropic's starting point | Notes |
|---|---|---|---|
| Claude Opus 5.5 | medium | medium, test low and high | Thinks more per turn at a given level than Opus 5 |
| Claude Sonnet 5.5 | high | medium for agents, low to medium for chat | Levels recalibrated from Sonnet 5 |
| Claude Opus 5 | high | high | Thinking can be disabled at high or below |
| Claude Sonnet 5 | high | high | medium comparable to Sonnet 4.6 at high |
Anthropic's own walkthrough of the Opus change is worth three minutes before you pick a level. In it, the Claude team fixes the same bug side by side on Opus 5 and Opus 5.5 in Claude Code, shows what each model does to plan usage limits, and covers "when medium effort is enough"; the video description adds that on Pro, Max and Team plans "your limits go 25% further".
The video is about Claude Code, not the raw API, but the lesson carries over: the default moved down a level because the model got more efficient, so the first experiment of any migration should be running your evals at medium and low before reaching for the setting you used on Opus 5.
Gateways add one more layer. OpenRouter's model metadata, read on October 6, lists the default effort for both 5.5 models as high and marks reasoning as mandatory - OpenRouter models API. Whatever a gateway does with an omitted parameter, it is not guaranteed to match Anthropic's own default. The fix is the same everywhere: set effort explicitly on every route, and change it mid-session with a per-message output_config (beta mid-conversation-output-config-2026-07-01) rather than the top-level field, because a top-level change invalidates the prompt cache.
Progress notes now arrive inside thinking blocks
On Opus 5 and Sonnet 5, the short notes a model writes between tool calls ("found the failing test, editing auth.py next") came back as text blocks, and many agent UIs streamed them as live status. On the 5.5 models, notes longer than a sentence or two come back as progress-update thinking blocks, at most one before each tool call, and under the default display: "omitted" their text is empty - progress updates. No request fails. The UI just goes silent for the length of a long agentic turn, and users assume it hung.
The fix has three parts. Set thinking.display to "updates" (beta header thinking-display-updates-2026-08-18) to receive a short summary of each note while reasoning stays hidden, then render every non-empty thinking block ahead of the tool_use it precedes. If the model may need to hand the user something verbatim mid-turn, give it a simple send-message tool, declared in the first request so the tools array never changes. And if turns still go quiet, Anthropic's documented trick is a turn-scoped system message after five silent tool steps reading "The user hasn't heard from you in a while - say in a few words what you're doing, then continue." which, in its testing on agentic coding tasks, "roughly halved the share of tasks with a long silent stretch, with no measurable change in cost".
A related Opus 5.5 behavior breaks unattended loops. On long multi-part tasks the model keeps the user updated, and some of those updates end the turn with text and stop_reason: "end_turn" rather than a tool call. A loop that treats any text-only end of turn as "task complete" stops halfway, which is the same failure we described for Claude Code's own loop in our guide to running Claude Code unattended. Anthropic's guidance is to keep a checklist the model updates, send a short continuation message naming open items when a turn ends without a stated blocker, and cap automatic continuations at two or three so a genuinely stuck run still ends - unattended agentic runs.
Refusals are HTTP 200, and some of them now cost money
Both 5.5 models run safety classifiers. A decline is not an error: it is a successful HTTP 200 with stop_reason: "refusal", an empty or partial content, and a stop_details object naming the category. Opus 5.5 adds a biology classifier that Opus 5 did not have; Sonnet 5.5 declines in five categories - Sonnet 5.5 migration guide. Monitoring built on error rates never sees any of this, which is why Anthropic tells you to instrument refusals as their own metric.
On September 24, Anthropic changed what a refusal costs. A refusal that arrives before any output is now billed when its category is bio, frontier_llm or reasoning_extraction, "the categories where Anthropic measures low volumes of false positives, as of September 2026", charged like any other request at the rates of the model that ran it. Pre-output refusals in other categories, or with a null category, stay free. Mid-stream refusals were always billed for the input and the output already streamed. Every refusal counts against rate limits either way.
| Category | What it covers | Billed before any output | Server-side fallback on Sonnet 5.5 |
|---|---|---|---|
cyber | Malware, exploit development; benign security work can trigger it | No | Retried on Sonnet 5 |
bio | Dangerous lab methods; beneficial life-sciences work can trigger it | Yes | Not retried |
frontier_llm | Helping build competing AI models | Yes | Retried on Sonnet 5 |
reasoning_extraction | Asking the model to write out its internal reasoning | Yes | Not retried |
general_harms | Other usage-policy areas | No | Not retried |
The reasoning_extraction row is the one most likely to hit ordinary production code, because it is triggered by prompts, not by subject matter. A <thinking> scratchpad section the model fills in, a reasoning or trace field in your JSON schema or tool input, "show your full reasoning" instructions, a running log of the model's private notes: all of these can be declined, now at full price, and server-side fallback will not retry them on either 5.5 model - Keep reasoning in thinking blocks. Audit your schemas and system prompts for those fields and replace them with a request for a short explanation, or with display: "summarized".
For legitimate declines, the fix is fallback, and Anthropic's own advice is to ship it from day one. The simplest form is server-side: fallbacks: "default" with the server-side-fallback-2026-07-01 header, which retries on the model Anthropic recommends for each category (for Opus 5.5, the launch post says "most cybersecurity tasks will be re-routed to Opus 4.8"). You can instead name up to three fallback models, each of which must appear in the requested model's allowed_fallback_models list on the Models API.
response = client.beta.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
output_config={"effort": "medium"},
fallbacks="default",
betas= ["server-side-fallback-2026-07-01"],
messages=messages,
)
served_by_fallback = any(
it.type == "fallback_message" for it in (response.usage.iterations or [])
) and response.stop_reason != "refusal"
if response.stop_reason == "refusal":
log_refusal(response.stop_details) # a metric, not an exception
Server-side fallback has limits worth writing on the whiteboard. It is a Claude API beta: it is rejected on the Batches API and unavailable on Bedrock, Google Cloud and Foundry, where the SDKs' refusal-fallback middleware does the same job client-side. It does not propagate into sub-agent calls made inside tool execution, so each sub-agent request needs its own. It applies sticky routing: after a fallback, later turns of the same conversation go straight to the fallback model for about an hour. And if you build the retry yourself, the fallback-credit-2026-07-01 header gives you a credit token (valid five minutes) that reprices the retry so you do not pay to cache the conversation twice - fallback credit.
Why this matters: these changes do not appear in error logs, so they surface as complaints and invoices weeks after the migration. How to apply it: set effort explicitly everywhere, render progress-update blocks, fix text-only end-of-turn handling in unattended loops, strip reasoning fields from schemas and prompts, turn on fallback for every request path, and add refusals and fallback-served responses as two separate counters on your dashboard.
7. What the Ecosystem Broke: Frameworks, Gateways and Claude Code
Most production code does not call the Messages API directly. It calls LangChain, the Vercel AI SDK, LiteLLM, Pydantic AI, Instructor or a gateway, and those libraries made the forced-tool-call decision on your behalf years ago. When the 5.5 models stopped accepting tool_choice: any, the breakage surfaced in the libraries first, which is why an application that never wrote the string tool_choice can still fail on its first request to claude-opus-5-5.
The failures come in two flavors, and the second one is worse. Loud failures are the 400s: a structured-output helper forces a tool and the API rejects it. Silent failures are requests that succeed while doing something different: a helper that quietly downgrades to an unforced tool and gets plain text back, or a wrapper that drops a disabled thinking setting so the model runs full adaptive thinking at full cost. The table summarizes what had shipped, and what was still open, as of October 6, 2026.
| Library | What broke | Status on Oct 6 |
|---|---|---|
| langchain-anthropic | with_structured_output forced the tool on Opus 5.5 | Fixed in 1.7.3 (Sep 22); Sonnet 5.5 in 1.7.5 (Sep 29); create_agent with ToolStrategy still open |
| Vercel AI SDK | toolChoice: "required" and jsonTool output returned 400 | Fixed in @ai-sdk/anthropic 4.0.60 (Sep 22) and 4.0.67 for Sonnet 5.5 |
| LiteLLM | tool_choice: "required" mapped to any and hit a 400 | v1.104.0 rejects locally unless drop_params downgrades to auto |
| Pydantic AI | OpenRouter route forced tool_choice='required' | Fixed in v2.52.0 (Sep 30) |
| Instructor | Bedrock TOOLS mode always forces the tool | Open; use MD_JSON mode or from_anthropic with auto |
The details behind those rows are where the migration risk sits. LangChain's fix for with_structured_output binds the tool unforced, warns you to switch to method='json_schema', and raises an OutputParserException when no tool call comes back, so code that never handled that exception now has a new failure path - LangChain PR 40766. But bind_tools(tool_choice=...), agents built with create_agent(response_format=ToolStrategy), and thinking=disabled still reach Opus 5.5 unmodified, per an open issue - LangChain issue 40777. The Vercel AI SDK takes the opposite approach and rewrites requests: it falls back to auto and to native output_config.format with a warning, and maps reasoning: 'none' to effort: 'low' - Vercel AI SDK PR 21286. Its Sonnet 5.5 release goes one step further and turns a disabled thinking setting into between_tools - @ai-sdk/anthropic changelog.
LiteLLM chose a third path: since v1.104.0, a tool_choice: "required" request to Opus 5.5 is rejected locally with an UnsupportedParamsError unless you set drop_params, which silently downgrades it to auto - LiteLLM PR 42489. Three related LiteLLM items were still open on October 6: the Bedrock response_format silent-text bug from section 3, a Sonnet 5.5 path where a dropped disabled setting leaves the model on full adaptive thinking, and schemas sent through the deprecated output_format field, which is ignored when effort is set. Instructor's Bedrock integration still forces toolChoice even when the caller passes auto - Instructor issue 2681. Pydantic AI's fix notes that its OpenRouter, Vercel, Copilot, Heroku, LiteLLM and Snowflake routes had all kept forcing until then - Pydantic AI issue 8675.
The lesson is not "upgrade your libraries", though you should. It is that a library fix changes semantics, and you need to know which semantics you got. Three libraries resolved the same API change in three different ways (unforced tool with an exception, silent rewrite to structured outputs, local rejection unless you opt into a downgrade). After upgrading, read the changelog entry, then test the specific code path: does it raise, warn, rewrite or degrade? Pin the version you tested, because the next minor release may resolve the open issues in yet another way.
Gateways pass the 400 through, and their metadata can mislead
Gateways add their own wrinkles. OpenRouter passes the upstream 400 straight through as "Provider returned error" with Anthropic's raw message, which is at least debuggable. Its model metadata, though, lists tool_choice and temperature among the accepted parameters for both 5.5 models with no caveat, and as section 6 noted, lists a default effort of high - OpenRouter models API. A gateway's parameter list tells you what it will forward, not what the model will accept. OpenRouter's own migration note also says it does not apply the preserved-thinking enforcement, so a harness that passes on OpenRouter can still fail on a new direct account - OpenRouter's Opus 5.5 migration note.
On Amazon Bedrock's Converse API, toolChoice: {any: {}} or a named tool returns a ValidationException carrying the same Anthropic text, which is how LlamaIndex's BedrockConverse integration and LiveKit's agents framework found the problem - LlamaIndex PR 23208. If you run a gateway in front of several labs, the cleanest fix is to translate an OpenAI-style tool_choice: "required" into auto plus a prompt instruction and a verification step at the gateway for these models, rather than letting each client discover the 400.
Claude Code hit the edges too
Even Anthropic's own harness needed several releases to settle. Claude Code 2.1.277 (September 18) added the AGENTS.md fallback: in a project with no CLAUDE.md, it now reads AGENTS.md instead, which matters if you keep one instruction file for several AI coders, as we covered in our AGENTS.md versus CLAUDE.md guide. Version 2.1.280 (September 22) added Opus 5.5 as the default Opus model, changed the default model on Pro and Team Standard plans from Sonnet to Opus, and stopped applying effort levels saved before per-model effort existed to newly released models - Claude Code changelog.
The fixes that followed read like a list of this guide's breaking changes. Version 2.1.281 stopped the settings UI from offering to turn thinking off on models that cannot. Version 2.1.282 fixed a failed turn with the error "Effort 'xhigh' isn't available with thinking turned off" after a safety-related model switch, which is exactly the fallback-plus-thinking-setting trap from section 2. Version 2.1.284 (September 28) added Sonnet 5.5 as the default Sonnet. And an open issue reports that 2.1.287's auto-mode classifier sends thinking.type: "disabled" on Bedrock, so Opus 5.5 and Sonnet 5.5 return 400 and block every Bash call, while 2.1.285 works - Claude Code issue 98962. Anthropic's own cookbook example for tool_choice returns a 400 on Fable 5.1 as well, per an open cookbook issue - claude-cookbooks issue 854.
There is a constructive reading of all this. If the team that designed the API needed four point releases to absorb it, the migration is genuinely hard, and nobody should feel bad about a staged rollout. And for integrations that sit on top of a harness rather than raw requests, most of the work is done by the harness: the docs list Claude Code, claude.ai, Managed Agents and the Agent SDK as already append-only. One concrete case is Founden, the platform publishing this guide, where desktop builds run through Claude Code on the founder's own Mac: as of October 6 its default model for Claude Code builds is Claude Opus 5.5, picked by the catalog rule that moves the default to the strongest model the coder's plan pays for, with no request bodies to rewrite because the harness owns the API calls.
Why this matters: your exposure to this migration is mostly determined by which libraries sit between your code and the API, and each resolved the change differently. How to apply it: inventory every LLM library and gateway in your stack, upgrade each to a version with explicit 5.5 handling, test what it now does on the forced-tool and thinking-off paths, and pin it.
8. Opus 5.5 or Sonnet 5.5: Where to Land
Once the code works on both models, the choice between them is a measurement problem, and the measurements disagree in instructive ways. Anthropic's launch table puts Sonnet 5.5 ahead of Opus 5.5 on Terminal-Bench 4.0, 70.6% against 66.4% (Opus at xhigh) - Anthropic's Sonnet 5.5 announcement. Artificial Analysis also measured Sonnet ahead, 64% against 60% - Artificial Analysis on Sonnet 5.5. The official leaderboard, running both at max effort inside Claude Code over 330 trials, puts Opus 5.5 first at 64.85% and Sonnet 5.5 at 61.82% - Terminal-Bench 4.0 leaderboard.
The disagreement is the finding. Harness, effort level and trial count move the ranking by more than the gap between the models, which means neither model wins by default and the only ranking that matters is the one on your own eval. The broader signals agree on that closeness: on the Artificial Analysis Intelligence Index, Opus 5.5 at max scores 58, which Artificial Analysis called "the highest score we have measured by several points", and Sonnet 5.5 scores 56 - Artificial Analysis on Opus 5.5. On LMArena's text leaderboard, Opus 5.5 sits fourth at 1504, statistically tied with Opus 4.6 - LMArena text leaderboard.
Effort, not the model, sets the cost and the latency
Where the two models clearly differ is in how cost and latency scale with effort. Artificial Analysis publishes per-effort measurements, and they show Sonnet 5.5 is dramatically cheaper and faster at the low end, while the advantage shrinks and then inverts at max effort, where Sonnet 5.5 used about 193,000 output tokens per task, "the highest token use we have measured". The chart below shows measured cost per task at three effort levels.
Latency follows the same curve. Artificial Analysis measured time to first token (which includes thinking) at 1.03 seconds for Sonnet 5.5 at low and 7.22 seconds at medium, against 8.29 and 23.54 seconds for Opus 5.5, with a median of 3.84 seconds for comparable reasoning models - Artificial Analysis, Sonnet 5.5 low. For a chat product or any route where a human watches a spinner, that gap decides the question before quality does. For a background agent that runs for twenty minutes, it barely matters.
Refusal behavior is a selection criterion now
The second real difference is how often each model declines. Vals.ai measured a 0.79% refusal rate and a 4.21% fallback rate for Opus 5.5, against 0.13% refusals for Sonnet 5.5 - Vals.ai on Opus 5.5. Averages hide the tail: on Vals's SRE benchmark, 217 of 262 Opus 5.5 tasks (82.82%) needed a fallback model, and counting those as failures drops the score from 33.59% to 5.34%. Sonnet 5.5 needed a fallback on 126 of 262 of the same tasks - Vals.ai on Sonnet 5.5.
Developers working near security felt this immediately. "Working in cybersecurity, I think I've been able to actually use Opus 5.5 maybe once or twice without triggering safeguards," one commenter wrote in the Opus 5.5 launch thread, which reached 1,806 points and 1,134 comments - Hacker News comment. Another complained about being charged "for thinking tokens that produce no result due to spurious refusals" - Hacker News comment, though under Anthropic's published rules a cyber refusal before any output is not billed; only bio, frontier_llm and reasoning_extraction are. For security tooling, operations automation or life-sciences work, run your eval with refusals counted as failures before choosing, and price in the fallback model.
Putting it together gives a usable default split. Sonnet 5.5 is the default for chat, extraction, classification, support agents and high-volume tool use at low or medium, where it is the cheapest and fastest frontier option per task. Opus 5.5 is the default for long-horizon coding, code review, multi-hour autonomous runs, analytical deliverables and computer use, where Anthropic and independent leaderboards agree it leads, and where its medium default already matches Opus 5 at max. Fable 5.1 is worth its $10/$50 only where your eval shows headroom above Opus 5.5, which Anthropic's own launch page suggests is rare, and our GPT-6 Astra versus Fable 5.1 comparison covers when that tier earns its price.
Anthropic's own short video on the two models is aimed at a different problem than the API changes, but it is the right thing to watch before your eval run: it covers "three habits that could be holding Opus 5.5 and Sonnet 5.5 back", starting from the point that newer models can take on bigger tasks than most people give them.
The habit worth carrying into the API work is the first one the video's description points at: briefing the model with the whole task rather than slicing it into small prompts. That is the same shift the API made when it removed forced tool calls, and it is why an eval built from your real, full-sized tasks predicts production far better than one built from toy prompts.
Why this matters: a wrong default here is a 2x cost error in one direction or a quality and latency error in the other. How to apply it: run the same 50 to 200 real tasks on both models at two effort levels each, count refusals as failures, record cost per completed task and time to first token, and pick per route, not per company. The previous-generation version of this decision, for teams still comparing, is in our Opus 5 versus Sonnet 5 guide.
9. The Cost Math of Migrating
Launch posts quote cost per token and cost per task. Your bill is neither: it is the product of your traffic shape and the model's behavior on it. Three shapes deserve separate math, because the 5.5 changes push them in different directions: short-output routes that used to run with thinking off, long agentic sessions that live on the prompt cache, and anything that triggers refusals.
Short-output routes: thinking is the whole bill
Take a classification route with a token profile we will state as an assumption, not a measurement: 2,000 input tokens per request, 1,500 of them a cached system prompt, and a 50-token answer. On Opus 5 with thinking disabled, that costs $4.50 per 1,000 requests at list prices. On Opus 5.5 with zero thinking it would cost $3.30, a 27% saving. But Opus 5.5 cannot disable thinking, and if it averages 200 thinking tokens on this route, the cost rises to $7.30 per 1,000, 62% more than before. The break-even is about 60 thinking tokens per request: below that Opus 5.5 is cheaper, above it more expensive.
The arithmetic explains why Sonnet 5.5 kept a thinking-off mode and Opus 5.5 did not. On Sonnet 5.5, between_tools gives the same route at $1.80 per 1,000 with no thinking at all, and even with 200 thinking tokens it stays below the old Opus 5 cost. For high-volume, short-output work (classification, routing, extraction, moderation), the migration that saves money is often Opus 5 to Sonnet 5.5, not Opus 5 to Opus 5.5. Measure thinking tokens per request at low effort on a sample of real traffic before you decide; usage.output_tokens includes them even when the thinking text is omitted. Our guide to setting the effort dial walks through that measurement.
Long agentic sessions: effort and the cache decide
For agents, the measured data tells a two-sided story. At its medium default, Opus 5.5 reaches an Artificial Analysis index of 51, the same as Opus 5 at max effort, for $1.34 per task against $5.86. At max effort, though, Opus 5.5 used about 119,000 output tokens per task against roughly 73,000 for Opus 5, 1.6 times as many, and ended up "level with Opus 5" on cost per task - Artificial Analysis on Opus 5.5. Sonnet 5.5 at max cost about $7.60 per task, roughly 50% more than Sonnet 5. Anthropic's "40% less to run" and "up to 30% less for most work" claims hold at moderate effort and invert at the top.
The prompt cache is the other half. Opus 5.5 bills cache reads at $0.20 per million tokens, 0.05x its input price, while a five-minute cache write costs $5. A miss is therefore 25 times the price of a hit on Opus 5.5, against 12.5 times on Opus 5 ($6.25 against $0.50). An agent that re-reads a 200,000-token cached prefix 50 times pays $2.00 for those reads on Opus 5.5 against $5.00 on Opus 5, but every invalidated prefix hurts relatively more. That is why the append-only discipline from section 4 is a cost measure as much as a correctness measure; our prompt caching guide covers the silent invalidators.
Routers deserve a line of their own here. A cost router that bounces a conversation between Opus 5.5 and Sonnet 5.5 pays twice: once for the cache it abandons (caches are per model) and once for the reasoning the new model cannot read (section 4). The routing savings in our model routing guide still hold, but on the 5.5 models they are best captured by routing per task or per session, not per turn.
Refusals: small per event, large per retry loop
The refusal billing change is cheap per event and expensive when code mishandles it. A pre-output refusal in a billed category charges the input tokens at the model's rate, so a refused request carrying a 50,000-token context costs $0.20 on Opus 5.5 and $0.10 on Sonnet 5.5. One engineer documented a billed $0.02 refusal on a biology lysis prompt that an open-weight model then answered for $0.05 - bede.im. The real risk is a retry loop that resends a refused request to the same model three times, which burns the cost three times and earns three refusals, and fallback that adds a billed refusal on top of the fallback request each time.
The defenses are simple. Never retry a refusal on the same model. Route it to fallback once, and on reasoning_extraction fix the prompt instead of retrying at all. Budget retries per request rather than per turn, since an agent and its sub-agents can each produce refusals. And if you pass model costs through to customers, add refusals to your metering so a classifier false positive does not become an unbilled cost; our guides to pricing an AI product against token costs and metered billing cover the pass-through mechanics.
Why this matters: the same migration can cut a bill by half or raise it by 60%, depending on traffic shape, effort and cache discipline. How to apply it: model each route's token profile separately, measure thinking tokens per request at your chosen effort, compare cost per completed task rather than per token, protect cache hit rates with append-only histories, and treat refusals as a metered event.
10. A Claude Opus 5.5 Migration Runbook You Can Ship This Week
Everything above collapses into a sequence. The order matters: find every affected request before changing any, isolate the model-specific shapes so the fix happens once, make history append-only before you trust any eval, and only then tune effort and roll out. Anthropic also ships an automated first pass: in Claude Code, running /claude-api migrate this project to claude-opus-5-5 applies the model ID swap and breaking parameter changes across a codebase and produces a checklist, after asking you to confirm the scope - Opus 5.5 migration guide. Treat its output as a draft to review, not a finished migration.
Step 1, inventory. Search the codebase and every installed LLM library for the request shapes that changed. The commands below catch the direct cases; repeat them inside node_modules or your Python environment for the libraries from section 7, and grep your prompt templates and JSON schemas for reasoning fields that invite reasoning_extraction declines.
rg -n "tool_choice" --glob '!node_modules' # forced tool use
rg -n "\"disabled\"|type: ?\"disabled\"|budget_tokens" # thinking off or budgets
rg -n "temperature|top_p|top_k" --glob '*.{py,ts,js}' # sampling params (400 if non-default)
rg -n "computer_20251124|computer_20250124|fine-grained-tool-streaming"
rg -n "content\ [0\]|\.content\ [0\]\.text" # positional block reads
rg -n "<thinking>|<reasoning>|\"reasoning\"|scratchpad" prompts/ schemas/
Step 2, one adapter. Model-specific request shapes should live in exactly one function, not at every call site. The adapter below takes a model ID and an intent and returns the request fields that model accepts, plus the name of any tool call the caller must now verify, so a router or fallback never sends between_tools to a model that rejects it, and a forced call never reaches a 5.5 model.
FORCED_TOOL_OK = {"claude-opus-5", "claude-sonnet-5", "claude-opus-4-8"}
THINKING_OFF = {"claude-opus-5": {"type": "disabled"},
"claude-sonnet-5": {"type": "disabled"},
"claude-sonnet-5-5": {"type": "between_tools"}} # Opus 5.5 has no off switch
def request_fields(model: str, *, thinking_off: bool, tool: str | None, effort: str):
"""Return (fields to send, tool name the caller must verify or None)."""
fields: dict = {"output_config": {"effort": effort}} # always explicit
if thinking_off and model in THINKING_OFF:
if model == "claude-sonnet-5-5" and effort in ("xhigh", "max"):
raise ValueError("between_tools needs effort high or below")
fields ["thinking"] = THINKING_OFF [model]
if tool is None:
return fields, None
if model in FORCED_TOOL_OK:
fields ["tool_choice"] = {"type": "tool", "name": tool}
return fields, None
fields ["tool_choice"] = {"type": "auto"} # 5.5 models: steer, then verify
return fields, tool
Step 3, fix and validate. Apply sections 2 through 5: delete disabled thinking or map it through the adapter, move extraction to structured outputs and actions to auto plus strict plus verification, migrate the computer toolset, re-pair advisors. Add schema validation in front of every tool handler, on every platform, because Bedrock gives you no strict mode and every platform can hand you a call that should not run. Our guide on why AI apps corrupt data covers that validation layer in depth.
Step 4, the append-only audit. Run the drop_block audit from section 4 on your real harness, log input_transformations on every turn in staging, and replace each history edit with its API feature. Do this before the eval, because an eval run on a harness that silently drops reasoning measures a degraded model. If you have an older account, set prefix_mismatch_behavior explicitly so staging behaves like a brand-new production account.
Step 5, the effort sweep. Run 50 to 200 real tasks per route at two or three effort levels, record quality, cost per completed task, time to first token at the median and 95th percentile, and stop reasons. Size max_tokens from what the runs actually used plus headroom, and stream anything above about 21,000 tokens. Keep the winning effort explicit in code, never implied by a default.
Step 6, refusals. Turn on fallback for every request path, including retries, background workers and sub-agents, and add two counters: refusals by category, and fallback-served responses. Alert on the gap between them, since a refusal that no fallback served is a user who got nothing.
Step 7, rollout. Shadow a slice of production traffic to the new model and compare outputs offline, then canary a small percentage, watching the counters from steps 5 and 6. Keep the previous model configured in the adapter until the canary is clean. Watch the calendar too: Claude Sonnet 4.5 retires on the Claude API on November 30, 2026, and Haiku 4.5 is listed as retiring no sooner than October 15, so teams on those models have a deadline rather than a choice - model deprecations.
Why this matters: each stage catches a failure the next stage cannot see, and the ones that skip stage 4 ship degraded quality that no eval flags. How to apply it: run the seven steps per service, keep the adapter as the only place model names and shapes live, and keep it after the migration, because the next model will change the surface again.
11. What These Changes Signal About Building on Frontier Models
Step back from the code and the September releases describe a direction. Every model Anthropic shipped that month carries the same constraints, introduced on its top tier first and then rolled down the line within four weeks. It would be surprising if the next Haiku did not follow, though Anthropic has not said so, and this guide treats it as an inference, not a fact. The design principle is consistent: the request no longer steers the model's decisions; it states the contract (output shape, effort, budget) and the model decides how to meet it.
That has a cost that is easy to miss. Signed, conversation-bound, and for Sonnet 5.5 account-bound reasoning means a conversation's most valuable state, the model's reasoning about your task, cannot move between models, accounts or providers. A fallback, a router, a migration to a new account, or a move to another lab all reset it. None of this is hidden; Anthropic documents it plainly. But it shifts where durable state should live. If your agent's progress exists only inside thinking blocks, you have handed the continuity of your product to the provider. If it lives in your own store (a task list, a summary, tool results you persist), a model switch is an inconvenience rather than an outage.
The rest of the market is moving just as fast, which reinforces the point. In September OpenAI launched GPT-6 Sol at $2/$10 and GPT-6 Luna at $0.10/$0.50, half the price of their GPT-5.6 predecessors - OpenAI's GPT-6 Sol and Luna announcement, then shipped GPT-6.1 Sol at the same price on September 29, and its deprecations page lists a long October 23 shutdown wave that includes o3-mini, o4-mini, gpt-4.1-nano and gpt-4-turbo - OpenAI deprecations. Google's Gemini 3.8 Flash is priced at $0.75/$3.75 through December 31, 2026, then doubles - Gemini API pricing. Every lab's API is a moving target, on its own schedule.
So the durable assets are not any one model's request shapes. They are the adapter that isolates them, the eval built from your real tasks that tells you when a new model is better, the state you keep outside the model, and a second model from another lab that you have actually qualified on that eval, so a pricing change, a deprecation or a classifier that suddenly declines your domain is a configuration change rather than an emergency. Teams that build that way absorbed September in a sprint. Teams that did not are still reading 400s.
Why this matters: the next breaking change is already on someone's roadmap, and the cost of each one depends on how much of your product is welded to a single model's surface. How to apply it: after this migration, keep the adapter, keep the eval current, persist agent state you own, and qualify one cross-lab alternative per critical route. Our playbook on orchestrating multiple agents shows how to structure that state for multi-agent systems.
12. Conclusion: A Decision Framework
The Claude 5.5 migration is four breaking changes on the surface and a change of philosophy underneath. Thinking is always on (Opus 5.5) or off only through between_tools (Sonnet 5.5). Forced tool calls are gone, replaced by output contracts and verification. Reasoning is bound to the model, the conversation and, for Sonnet 5.5, the account. Computer use moved to a new toolset. Around those sit the changes that fail no request: the medium effort default on Opus 5.5, progress notes inside thinking blocks, and billed refusals in three categories since September 24.
The decision is easier than the mechanics, because most of it follows from your traffic shape rather than from benchmark tables. A route's latency budget, output length, domain and platform narrow the choice to one or two options before any eval runs, and the eval then settles the effort level. Use the framework below as the starting point, then let your own eval overrule it wherever the two disagree, since every benchmark in this guide was run on someone else's tasks.
| Your situation | Land on | Starting settings |
|---|---|---|
| High-volume, short-output, latency-sensitive routes | Claude Sonnet 5.5 | low effort or between_tools, structured outputs |
| Long-horizon coding, review and autonomous runs | Claude Opus 5.5 | medium, raised only where evals show a gain |
| Security, operations or life-sciences workloads | Either, after a refusal-aware eval | Fallback on every path from day one |
| Mid-launch, no capacity to migrate | Stay on Opus 5 or Sonnet 5 for one cycle | Inventory and adapter built now |
| Amazon Bedrock deployments | Either 5.5 model | No strict tools: validate every tool input in code |
What ties those together is the same rule: migrate to contracts, not to a model. Put request shapes behind one adapter, keep history append-only, keep state you own, set effort explicitly, and measure cost per completed task. Done that way, this migration is a week of work for most teams, and the next one will be a day. If your product reaches Claude through a harness such as Claude Code or the Agent SDK rather than raw requests, much of the conversation-binding work is already handled, which is one more reason to let a harness own the API calls where it fits.
This guide reflects the Claude API, pricing and third-party benchmarks as of October 6, 2026. Model behavior, prices, refusal billing categories and library versions change frequently, so verify current details in Anthropic's documentation before migrating production traffic.