Guide

Claude Sonnet 5.5, explained: the cheap model that keeps embarrassing Anthropic's own flagship

Anthropic shipped Claude Sonnet 5.5 on 28 September 2026, six days after Claude Opus 5.5 rather than before it, at the same $2 input / $10 output per million tokens as the Sonnet it replaces. On two of the benchmarks that matter most for everyday work, it lands within two points of Opus 5.5, a model that costs twice as much. The catch competitors keep missing: at Anthropic's own highest reasoning-effort setting, Sonnet 5.5 burns so many more tokens per task that the cost advantage can flip entirely.

به‌روزرسانی‌شده 9 Oct 202610 min read
Quick answer

Claude Sonnet 5.5 launched 28 September 2026, the second release in Anthropic's "5.5" refresh, six days after Opus 5.5 (22 September). Pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output tokens. The one price that did change is easy to miss: cached-prompt reads dropped to $0.10 per million tokens, half Sonnet 5's old $0.20 rate. Most third-party pricing trackers we checked still list $0.20 as current.

On knowledge-work and computer-use benchmarks, it's a near-tie with Opus 5.5 at half the price. Sonnet 5.5 scores 1844 on GDPval-AA v2.1 against Opus's 1846, and 80.1% on OSWorld 2.1 against Opus's 81.8%, gaps of one to two points despite the 2x price difference.

Anthropic's own vendor number and an independent lab's re-test disagree. Anthropic reports 70.6% on Terminal-Bench 4.0; Artificial Analysis's independent re-run of the same benchmark got 64%, still ahead of Opus 5.5's 60% on AA's harness, but a real gap between the marketing number and reproduced results worth knowing before you plan around it.

The "up to 30% cheaper" claim doesn't hold at every setting. At maximum reasoning effort, Sonnet 5.5 generates roughly 193,000 output tokens per task on average, the highest Artificial Analysis has measured for any model, which pushes its per-task cost to about $7.60, higher than Sonnet 5 and, on one measurement, higher than Opus 5.5 reaching the same quality bar with fewer, more expensive tokens.

Anthropic's API pricing page showing Sonnet 5.5's prompt caching rates, $0.10/MTok read and $2.50/MTok write, alongside Opus 5.5 and Fable 5.1 cards on the same page.
Anthropic's API pricing tab, confirming Sonnet 5.5's $0.10 cache-read rate directly. anthropic.com, October 2026.

What shipped, and in what order

Our own Opus 5.5 guide stated the opposite order until we corrected it while researching this piece. Anthropic's own model documentation lists Opus 5.5's release as 22 September 2026 and Sonnet 5.5's as 28 September 2026, six days later, not earlier. The order matters for how to read the comparison: Sonnet 5.5 is the follow-up act that had Opus 5.5's benchmark numbers to beat, not the warm-up.

The API model ID is claude-sonnet-5-5, available through the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. Context window stays at 1 million tokens, with a 128K-token output limit on standard requests, extendable to 300K tokens on the Batch API behind an output-300k-2026-03-24 beta header. Knowledge cutoff is June 2026 for both training and reliable-recall purposes. One behavioral default changed from Sonnet 5: Sonnet 5.5 reasons by default at high effort (Opus 5.5 defaults to medium), and its lowest manual thinking setting is now between_tools rather than the old disabled option, a breaking change if your integration explicitly set thinking off.

The pricing catch: cache reads got cheaper, and most trackers missed it

Input and output pricing carried over unchanged from Sonnet 5: $2 per million input tokens, $10 per million output. Prompt caching is where something actually moved. Anthropic's documentation states Sonnet 5.5 keeps the same prices as Sonnet 5, "except for prompt cache reads, which cost $0.10 USD per million tokens, half the Claude Sonnet 5 rate." Cache writes stay at $2.50/million for a 5-minute cache and $4/million for a 1-hour cache. The Batch API still applies a 50% discount across input and output, effectively $1/$5 per million tokens for asynchronous workloads.

We checked several pricing round-ups published after the Sonnet 5.5 launch, and more than one, including eesel.ai's pricing guide, still lists cache reads at $0.20, the old Sonnet 5 figure, rather than the corrected $0.10. If you're budgeting a high-cache-hit-rate workload, the gap compounds: a workload reading mostly from cache costs roughly half what an outdated chart would suggest. Check any pricing claim against Anthropic's own docs directly rather than a secondary table, including this one, before committing a budget to it.

Benchmarks: a near-tie with Opus, with one vendor-vs-independent gap

Benchmark                    Sonnet 5.5   Opus 5.5   Sonnet 5
----------------------------------------------------------------
GDPval-AA v2.1 (knowledge work)    1844         1846        1449
OSWorld 2.1 (partial credit)       80.1%        81.8%       57.0%
OSWorld 2.1 (strict pass)          43.5%        48.7%       25.6%
Terminal-Bench 4.0 (Anthropic)     70.6%        66.4%       10.3%
Terminal-Bench 4.0 (AA re-test)      64%          60%          -
FrontierCode 1.1 (xhigh effort)    52.1%        54.4%       42.4%
CursorBench 4.0                    55.5%        57.8%       34.1%
Humanity's Last Exam (w/ tools)    64.5%        67.7%       54.9%

AA = Artificial Analysis, an independent benchmark lab.
Terminal-Bench shows the largest vendor-vs-independent gap of
any figure checked for this article.

On GDPval-AA, a knowledge-work benchmark, and OSWorld, a test of operating a real computer through a GUI, Sonnet 5.5 lands within one to two points of Opus 5.5 despite costing half as much. That's the headline most coverage leads with, and it's accurate. What most coverage skips is the OSWorld split: the 80.1% figure is partial credit, rewarding progress toward a task even if it isn't fully completed. The strict pass rate, full task completion only, is 43.5% for Sonnet 5.5 against 48.7% for Opus 5.5, a gap that widens from 1.7 points on partial credit to 5.2 points on strict pass, a meaningfully different number to plan around if your use case needs a task finished rather than mostly finished.

The Terminal-Bench gap is worth flagging on its own. Anthropic's launch materials report 70.6% for Sonnet 5.5. Artificial Analysis, an independent lab that re-runs benchmarks on its own harness rather than taking vendor numbers at face value, measured 64%, a six-point drop from the vendor figure, though still ahead of Opus 5.5's 60% on the same harness. Neither number is wrong exactly; they reflect different test conditions. But a six-point gap between what a company reports and what an independent party reproduces is the kind of detail that belongs in any serious comparison, and it was largely absent from the competitor pages we checked.

One Hacker News thread discussing the launch (884 points, 613 comments) raised a related wrinkle: part of Sonnet 5.5's lead over Opus 5.5 on some agentic benchmarks may reflect Opus's safety-routing behavior rather than a pure capability gap. One commenter cited Opus 5.5 falling back to a weaker model in roughly 10% of trials due to safeguard triggers, well above Sonnet's fallback rate on the same tests. That's a community observation, not an Anthropic-confirmed figure, but it's a plausible explanation for why a cheaper model keeps nearly matching or edging out a pricier one on specific tests.

The real cost story: token verbosity at max effort

Anthropic markets Sonnet 5.5 as delivering tasks at up to 30% lower cost than Sonnet 5. That claim is real at Anthropic's chosen task mix and effort level, but it doesn't hold universally, and the gap is the single most useful fact in this article for anyone actually paying the bill. Artificial Analysis measured Sonnet 5.5 generating an average of roughly 193,000 output tokens per task at maximum reasoning effort, the highest figure AA has recorded for any model it tracks, about 60% more than Opus 5.5 or Sonnet 5 at the same max-effort setting, and roughly seven times GPT-6 Astra's verbosity on comparable tasks.

That verbosity shows up directly in cost per task: AA measured Sonnet 5.5 at roughly $7.60 per task at max effort, about 50% higher than Sonnet 5's cost at the same setting. On one AA comparison, Opus 5.5 reached Sonnet 5.5's peak quality score for about $3.46 per task, needing fewer, more expensive tokens rather than many cheaper ones to get there, cheaper in total than Sonnet 5.5's own max-effort run. The practical lesson: Sonnet 5.5's price advantage is real at its default high-effort setting for well-scoped tasks, and can disappear or reverse if you push it to max effort on something genuinely hard, where Opus 5.5's fewer, pricier tokens sometimes finish the job for less.

Decision diagram: well-scoped, repeatable tasks route to Claude Sonnet 5.5 at $2/$10 per million tokens; ambiguous or high-stakes tasks route to Claude Opus 5.5 at $4/$20 per million tokens.
Anthropic's own guidance still defaults unsure users to Opus 5.5, stepping down to Sonnet once a task proves well-scoped.

Sonnet 5.5 vs Opus 5.5: what Anthropic itself tells you to buy

Anthropic doesn't position Sonnet 5.5 as a replacement for Opus 5.5. Its own announcement calls Sonnet 5.5 "a faster, lower-cost complement to Claude Opus 5.5," reserving Opus for "complex work requiring careful judgment." The documentation's model-comparison table is blunter still: it tells undecided users to start with Opus 5.5 for most workloads, and treats Sonnet as the model you step down to once you already know a task is well-scoped, not the default choice.

There's a genuine accuracy tradeoff behind that guidance that benchmark tables don't show directly. On Artificial Analysis's Omniscience test, a measure of factual accuracy, Opus 5.5 scores 66% against Sonnet 5.5's 54%, Opus simply knows more facts correctly. But on hallucination rate, Sonnet 5.5 is actually lower (better) at 47% against Opus's 59%: when Sonnet 5.5 doesn't know something, it's more likely to say so rather than guess, even though it has less to work with overall. For research synthesis specifically, a lower hallucination rate on what it does attempt can matter more than a higher raw-knowledge ceiling, depending on whether you'd rather catch a gap or catch a wrong answer stated with confidence.

The complaint dominating developer discussion: safeguards blocking ordinary work

The loudest recurring theme in that Hacker News thread wasn't price or benchmarks, it was Anthropic's cyber-safeguard system flagging legitimate work as risky. Developers reported getting blocked or routed to a restricted mode on tasks like ESP32 Bluetooth firmware, old C code review, write-ahead-log implementations, and authorized bug-bounty research. One commenter's summary, "I do not want a nanny tool," was echoed repeatedly through the thread, alongside complaints that the underlying "Cyber Verification Program" produces false positives without a clear appeal path.

A second complaint cluster centered on usage limits rather than safety routing: multiple users reported exhausting Max-plan quotas well before the billing period renewed, one citing an Opus 5.5 20x plan used up in four of seven days, and a separate account of a team hitting a $10,000 monthly token budget in under a day of heavy agentic use. Neither complaint is specific to Sonnet 5.5's capabilities, but both affect the practical cost of adopting it inside an agentic workflow, and neither shows up in a benchmark chart.

One independently run comparison gives a more concrete sense of real-world reliability. Code-review platform CodeRabbit tested Sonnet 5.5 against Sonnet 5 on 13 pull requests with known, previously catalogued bugs: Sonnet 5.5 caught 6 of 13 against Sonnet 5's 4 of 13, at roughly 60% lower cost per review and about half the latency (a mean of 5 minutes 27 seconds versus 9 minutes 55 seconds). Opus 5.5 still caught more, 8 to 10 of the same 13 bugs depending on the run. That matches the pattern Anthropic's own positioning describes: Sonnet 5.5 is a real upgrade over Sonnet 5 at the same price, and Opus 5.5 still catches more on the cases that matter most.

Sonnet 5.5 against the field: GPT-6 Sol and the Pareto problem

OpenAI's GPT-6 Sol is priced identically to Sonnet 5.5 at $2/$10 per million tokens, which makes the two a direct head-to-head on cost. Sonnet 5.5 leads clearly on GDPval-AA (1844 vs. 1487) and on AA-Briefcase (1811 vs. 1483), and edges ahead on FrontierCode (52.1% vs. 49.3%). But Artificial Analysis notes Sonnet 5.5 sits off the intelligence-per-output-token frontier at matched price, because of the same verbosity problem described above: GPT-6 Sol delivers more measured capability per output token at a comparable price point, even where Sonnet 5.5 wins on raw benchmark scores. Which one actually costs less in production depends heavily on how verbose your specific workload lets a model get, not just the headline benchmark table.

Who should actually use Sonnet 5.5

  • Use Sonnet 5.5 if your workload is well-scoped and repeatable: routine coding, bug fixes with a clear reproduction case, or a subagent running under an Opus-level orchestrator. This is where its near-Opus benchmark scores at half the price actually pay off.
  • Use Sonnet 5.5 if you were already on Sonnet 5 and haven't changed anything. The upgrade is effectively free on the sonnet model alias, and independent testing (CodeRabbit's PR review comparison, a separate hands-on editing-agent test) shows a real quality jump at the same price.
  • Escalate to Opus 5.5 if the task is ambiguous, spans an unfamiliar codebase, or a single wrong assumption would be costly. Opus's edge over Sonnet widens on exactly these harder, less-scoped cases.
  • Don't assume "up to 30% cheaper" at max effort. If you're pushing Sonnet 5.5 to its highest reasoning setting on genuinely hard problems, check your actual token spend. Independent measurement found it can cost more than Sonnet 5, and sometimes more than Opus 5.5 reaching the same result.
  • Budget cache-heavy workloads using $0.10/million for reads, not the $0.20 figure several pricing trackers still publish.

People also ask

When did Claude Sonnet 5.5 launch, and before or after Opus 5.5?

28 September 2026, six days after Claude Opus 5.5 (22 September 2026), not before it. An earlier version of our own Opus 5.5 guide had the order reversed; we corrected it after verifying against Anthropic's own documentation.

How much does Claude Sonnet 5.5 cost?

$2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5. Cached-prompt reads cost $0.10 per million tokens, half Sonnet 5's old $0.20 rate, a change several third-party pricing trackers still haven't updated. The Batch API applies a 50% discount on top of standard pricing.

Is Claude Sonnet 5.5 as good as Opus 5.5?

On GDPval-AA (knowledge work) and OSWorld (computer use), Sonnet 5.5 scores within one to two points of Opus 5.5 despite costing half as much. Opus 5.5 pulls ahead on harder, less-scoped benchmarks and on factual accuracy (66% vs. 54% on AA-Omniscience), while Sonnet 5.5 actually hallucinates less when it doesn't know an answer (47% vs. 59%).

Is Claude Sonnet 5.5 really up to 30% cheaper than Sonnet 5?

At Anthropic's chosen task mix and default effort level, yes. At maximum reasoning effort, independent testing found Sonnet 5.5 generates far more output tokens per task than Sonnet 5, pushing its cost per task roughly 50% higher than Sonnet 5 at that setting, and sometimes higher than Opus 5.5 reaching the same quality with fewer tokens.

Why do people complain about Claude's cyber safeguards on Sonnet and Opus 5.5?

Developers have reported legitimate work, firmware code, old C review, authorized security research, getting flagged or restricted by Anthropic's Cyber Verification Program. The complaint isn't about a specific wrong answer; it's about false positives interrupting ordinary development work without a clear path to resolve them.

notebooklm-to-pdf.comهمه راهنماها →

NotebookLM خود را با یک کلیک خروجی بگیرید

افزونه رایگان Chrome. PDF، Word و Markdown. روی دستگاه شما رندر می‌شود — چیزی آپلود نمی‌شود.

ادامه مطالعه