Updated 3 hours ago
AI News
Claude Sonnet 5.5 halves cache-read prices. Migration can still change your bill
Input and output token prices match Sonnet 5, while cache reads now cost half as much. Effort defaults, thinking behavior and unsupported API settings still make migration more consequential than a model‑ID swap.
Input and output stayed flat; cache reads did not
Anthropic released Claude Sonnet 5.5 on September 28 with the same base input and output prices as Sonnet 5: $2 per million input tokens and $10 per million output tokens. Anthropic's current [pricing table](https://platform.claude.com/docs/en/about‑claude/pricing) lists a real discount elsewhere: cache reads cost $0.10 per million tokens on Sonnet 5.5, half Sonnet 5's $0.20 rate. Neither change is evidence that the same uncached workload will produce the same bill.
Anthropic says Sonnet 5.5 generates output more than 30% faster and costs up to 30% less per task because it often uses fewer tokens to finish work. Those are workload‑level claims from the model maker, not a cut to the $2/$10 base rates. An independent launch evaluation from [Artificial Analysis](https://artificialanalysis.ai/articles/claude‑sonnet‑5‑5) found the opposite at the maximum effort setting: roughly 193,000 output tokens and $7.60 per Intelligence Index task, about 50% more than Sonnet 5 in that suite. The evaluator described it as the highest output‑token use it had measured.
Both findings can be true. A lower or middle effort setting can finish many well‑scoped tasks efficiently, while the maximum setting spends far more tokens pursuing a higher score. Anthropic's own launch charts make the same general point: cost and score move with effort, and Sonnet 5.5 complements Opus most clearly at lower settings. Teams should treat “up to 30% less” as a hypothesis to test on a pinned workload, not as a new base input/output price.
Medium and High are different starting lines
The default effort depends on where the model is used. In Claude apps and Claude Code, Sonnet 5.5 defaults to Medium. On the Claude Platform API, it defaults to High. Anthropic recommends Medium for well‑specified agentic tasks and High for harder or longer work, while latency‑sensitive chat may start at Medium or Low.
That split can distort a casual evaluation. A developer who compares the app with an API integration without pinning effort may be comparing two reasoning budgets as well as two interfaces. A production migration that simply swaps the model name and accepts the API default can also change latency and token use even when the prompt stays identical.
The safer sequence is to freeze the task set, set `output_config.effort` explicitly and run at least two plausible levels. Record strict task completion, output tokens, wall‑clock latency, retries and human repair time. The public [Claude Sonnet 5.5 model page](https://opentools.ai/llms/claude‑sonnet‑5‑5) is a useful place to compare its stable model facts; use Anthropic's current pricing table for cache‑read rates, and test the launch‑specific decision on the organization's own workload and chosen effort.
Several old request patterns now return errors
Anthropic's [migration guide](https://platform.claude.com/docs/en/models/sonnet‑5‑5/migration‑guide) lists changes that can turn a model swap into a 400 error. Sonnet 5's `thinking: {"type": "disabled"}` is not supported. To keep up‑front thinking off, Sonnet 5.5 uses `between_tools`, and only at Low, Medium or High effort. Xhigh and Max require adaptive thinking. `between_tools` also prevents changing effort in the middle of a conversation.
Forced tool choice changes too. Requests that use a tool choice of `any` or a named `tool` are rejected. Anthropic directs Claude API users to `auto` with strict tool schemas where supported, and to express the desired tool behavior in the prompt. Amazon Bedrock has a separate limitation: structured outputs, including strict tool use, are not available for Sonnet 5.5 there, so applications must validate tool inputs themselves.
The response parser may need work even when the request succeeds. Thinking runs by default when the `thinking` field is omitted, and a response can begin with a thinking block. Code that assumes `content[0].text` can therefore break. Tool loops must return thinking blocks unchanged, and `max_tokens` must cover thinking plus answer text. Other starting models add more changes: fixed thinking budgets, non‑default sampling parameters and assistant prefills can be rejected. The official checklist should be applied from the exact model and platform being replaced.
The benchmark jump needs its configuration footnotes
Anthropic reports 70.6% for Sonnet 5.5 and 10.3% for Sonnet 5 on Terminal‑Bench 4.0. That is a striking launch number, but it should not be carried into a capacity plan as a universal completion rate. The company describes the evaluation in an effort‑versus‑cost framework, and its page says benchmark scores capture only one facet of capability. Artificial Analysis reports a different Terminal‑Bench result—64% at maximum effort—because it used its own evaluation configuration.
The independent evaluator also found Sonnet 5.5 close to Opus 5.5 on its aggregate Intelligence Index while using substantially more output tokens. On factual knowledge, it remained behind Opus in that suite. These results support a narrower conclusion: Sonnet 5.5 can reach very strong agentic performance, but the effort setting and harness determine how much it costs to get there.
There is another reason to wait for a clean rerun before treating every launch comparison as settled. Anthropic says Artificial Analysis tested a pre‑release deployment with a structured‑output bug that could degrade some responses. The evaluator says it will rerun the affected tests. Anthropic expects the impact to be small or to understate performance, but the direction and size have not yet been independently measured. Artificial Analysis still labels those evaluations as pre‑release and says the affected tests will be rerun. The article's cost figures therefore describe the retained launch run, not a forecast of any rerun.
A migration check should include safeguards and fallback
Sonnet 5.5 is the first Sonnet release to ship with cyber safeguards similar to Anthropic's more capable models. Anthropic says ordinary software development should remain unaffected, but higher‑risk cyber requests can be refused or, in some Claude API configurations, retried on Sonnet 5 through server‑side fallback. Biology and other policy categories have different fallback behavior. An application that treats every refusal as an infrastructure failure can misroute or repeatedly retry a policy decision.
Before changing production traffic, teams should handle the documented `refusal` stop reason, log any fallback model and compare it against the expected policy. Security teams should include benign defensive cases that sit near the boundary, because a model that succeeds on ordinary code generation may behave differently on vulnerability work. They should also keep conversations append‑only when reusing signed thinking blocks; edits to earlier history can produce an error on accounts where that integrity check is enforced.
A useful acceptance gate is more concrete than “the new model scored higher.” Pin the platform, model ID and effort. Replace unsupported thinking and tool settings. Read response blocks by type. Run the existing regression set at two effort levels. Measure completed tasks and total intervention cost, not only token price. Exercise refusals and fallback. Only then move traffic. Sonnet 5.5 may be faster and cheaper for many routine jobs, but the public evidence does not support assuming that outcome for an unmeasured maximum‑effort agent.
*Figure: Claude Sonnet 5.5 migration decision points. Sources: [Anthropic pricing](https://platform.claude.com/docs/en/about‑claude/pricing), [Anthropic migration guide](https://platform.claude.com/docs/en/models/sonnet‑5‑5/migration‑guide), [Anthropic launch announcement](https://www.anthropic.com/claude‑sonnet‑5‑5) and [Artificial Analysis](https://artificialanalysis.ai/articles/claude‑sonnet‑5‑5), accessed October 8, 2026. Figure by OpenTools Team; no third‑party expressive material used.*
Related News
Oct 8, 2026
Anthropic’s nine influence cases show distribution is not persuasion
A case-by-case reading of Anthropic’s September threat report separates AI-generated output, verified audience reach and evidence of real-world effects.
AnthropicAI influence operationsAI safety
Sep 29, 2026
GPT-6.1 Sol launches at GPT-6 Sol prices, with cached input cut in half
OpenAI says GPT-6.1 Sol approaches Astra on agentic work without raising standard token prices. The launch data is promising, but its benchmark settings still matter.
OpenAIGPT-6.1 SolAI models
Sep 28, 2026
Claude’s nine-loop result is public, but “verified” has four different meanings
Claude’s nine-loop amplitude has unusually detailed checks. The public record shows where exact agreement ends and independent reproduction is still missing.
AnthropicClaudeAI science