By 19:30 UTC on release day, Claude Sonnet 5.5 had drawn 569 Hacker News points and 396 comments. The more immediate number for developers is five: five API settings can turn an apparently routine model swap into a 400 error. Anthropic's September 28 release pairs a claimed 30% speed gain with lower token use, but getting that gain may require changes to thinking controls, response parsing, tool selection, and effort settings.
That distinction matters because Anthropic has not cut Sonnet's list price. Sonnet 5.5 still costs $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens. The company's claim of up to 30% lower cost applies to a completed task, based on the model taking fewer tokens and fewer tool steps to finish work. Teams that budget only from the rate card could miss both halves of the release: the possible savings and the migration cost.
The discount depends on finishing in fewer steps
Anthropic positions Sonnet 5.5 as the everyday counterpart to Opus 5.5. It is aimed at bounded work such as fixing bugs and producing documents, while Opus remains the company's choice for open-ended work that needs sustained judgment. The model is available through Anthropic's API and on Amazon Web Services, Google Cloud, and Microsoft Azure under the ID claude-sonnet-5-5. Claude Code and Anthropic's apps default to medium effort, while the API platform defaults to high effort, according to the launch details.
The price claim rests on work done per request, rather than a cheaper token. Anthropic says Sonnet 5.5 generates output at least 30% faster than Sonnet 5. In the company's tests, low or medium effort beat Sonnet 5's best result on several evaluations for roughly one tenth of the cost per task. That is an appealing ratio for agent loops, where every extra shell command, search, or correction adds latency and billed output. It is still a vendor measurement, and the page does not turn 'up to 30%' into a universal discount for every prompt mix.
Customer tests published by Anthropic point in the same direction, with different magnitudes. Slack reported about 14% fewer output tokens in its offline Slackbot evaluations. Lovable measured about one third fewer tool calls and roughly half as many shell runs. Balyasny Asset Management said Sonnet 5.5 used about 121,000 tokens per answer across 2,441 finance tasks, compared with 497,000 for Sonnet 5. These figures came from early testers quoted by the model maker, not a common independent test harness, so they are useful migration clues rather than a clean cross-company ranking.
Five familiar settings can now fail
The official migration guide lists five settings that the basic Sonnet 5.5 request must leave out: explicit thinking budgets, non-default sampling parameters, assistant prefill, forced tool choice, and thinking: {"type": "disabled"}. Depending on the setting, old code can receive a 400 response instead of a subtly different answer. This makes a production rollout an API migration, even when the model name is the only planned product change.
Thinking is the sharpest edge. A request without a thinking field now uses adaptive thinking. Applications that want no thinking before the first tool call must use between_tools, not disabled. That mode works at low, medium, and high effort, but Anthropic says it returns a 400 error at xhigh or max. It also rejects extra display, budget_tokens, or block_binding fields, and its effort level cannot change partway through a conversation.
A minimal low-thinking request therefore looks like this:
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
thinking={"type": "between_tools"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": task}],
)
The guide also warns that responses can begin with a thinking block. Code that assumes response.content[0].text may break, even if the request succeeds. Tool loops must return thinking blocks unchanged, including empty ones. Since max_tokens covers thinking and visible text, and thinking tokens are billed as output, a team should remeasure token ceilings and cost instead of copying its Sonnet 5 limits.
Forced tool calls require another decision. Anthropic directs developers to use automatic tool selection with strict tools, or automatic selection alone on Amazon Bedrock. The migration checklist also calls for append-only conversations and explicit handling of refusals and fallbacks. Those are observable behavior changes around an agent loop, not cosmetic parameter renames.
The benchmark jump needs its footnotes
Anthropic reports a 70.6% score for Sonnet 5.5 on Terminal-Bench 4.0, up from 10.3% for Sonnet 5. On CursorBench 4.0, which uses ambiguous multi-file tasks drawn from Cursor sessions, the new model scored 55.5%, versus 34.1% for Sonnet 5 and 57.8% for Opus 5.5. FrontierCode produced a less tidy result: Sonnet 5.5 scored 52.1% at xhigh effort but fell to 46.2% at max.
That reversal is more informative than another top-line win. Anthropic says FrontierCode penalizes changes that would not be merged without human edits. At maximum effort, Sonnet 5.5 invoked Claude Code's review skill more often, splitting work among subagents. In two cases examined by Cognition, that choice caused a timeout or edits beyond the task's scope. More computation produced a worse measured outcome because the agent did extra work.
The example also limits what the published scores establish. They show that the tested model and agent setup performed better on those task sets at particular effort levels. They do not show that maximum effort is safest, cheapest, or most accurate for an arbitrary repository. Anthropic itself says Opus 5.5 remains stronger on complex work requiring judgment, despite Sonnet approaching it on several charts.
Faster coding brings stronger cyber controls
Sonnet 5.5 is the first Sonnet release to ship with the cyber controls Anthropic previously reserved for its most capable models. The company says higher-risk cybersecurity requests may visibly fall back to Sonnet 5, while ordinary debugging and defensive work remain available. It also added classifiers intended to deter large-scale reasoning extraction and expanded preserved thinking so reasoning stays tied to the account that created it. Developers who move a conversation between accounts, including during a Claude Code session, may encounter that boundary.
Anthropic says its automated behavioral audit covered about 1,850 scenarios and found that Sonnet 5.5 matched or improved on Sonnet 5 across most measures. The company also states that it found no evidence of the model pursuing goals against the user's intention. Its release page immediately narrows that claim: no evaluation set catches every failure, and untested tendencies may remain. That caveat should travel with the score when teams assess long-running agents that can edit files or call external services.
Measure completed work before switching traffic
A sensible evaluation unit for Sonnet 5.5 is the finished task. Record whether the change passed, how long it took, what the full trace cost, and which tools failed. Then compare those results at low, medium, and high effort on work drawn from your own queue. Anthropic's data suggests that a lower effort setting can beat Sonnet 5 while spending much less, while its FrontierCode footnote shows how a larger reasoning budget can invite extra actions and a lower score.
The next evidence should come from production traces after teams finish the migration, especially workloads with long tool loops and strict completion checks. Watch whether the promised reduction in tokens survives real prompts, and whether fallback or preserved-thinking rules interrupt existing sessions. If Sonnet 5.5 earns its discount, it will appear in fewer retries and a smaller bill for accepted work, not in the unchanged per-token price.