Anthropic released Claude Sonnet 5.5 this evening, and the headline number is not a price cut, it's a benchmark jump. On Terminal-Bench 4.0, the test Anthropic now leans on to show how a model handles long, unsupervised agent work, Sonnet 5.5 scores 70.6%. Sonnet 5 scored 10.3% on the same test. That's not a typo: Anthropic rebuilt the mid-tier model specifically for extended agentic loops, and it now beats Anthropic's own flagship Opus 5.5, which scores 66.4% on Terminal-Bench 4.0, while charging half as much per token.

  • Sonnet 5.5 launched September 28, 2026, the second release in the Claude 5.5 family after Opus 5.5 shipped six days earlier.
  • Terminal-Bench 4.0 jumped from 10.3% (Sonnet 5) to 70.6%, ahead of Opus 5.5's 66.4% on the identical benchmark.
  • List pricing is unchanged at $2 per million input tokens and $10 per million output tokens. Anthropic's "up to 30% cheaper" claim is about finishing a task in fewer billed tokens, not a lower rate.
  • It's live now on the Claude Platform, AWS, Google Cloud and Microsoft Azure as claude-sonnet-5-5, with early adopters including Epic Games, Zendesk, Unity and SpaceXAI.
Terminal-Bench 4.0: Sonnet 5 vs Opus 5.5 vs Sonnet 5.5 A bar chart showing Claude Sonnet 5 scoring 10.3 percent on Terminal-Bench 4.0, Claude Opus 5.5 scoring 66.4 percent, and the new Claude Sonnet 5.5 scoring 70.6 percent, ahead of Anthropic's own flagship model on this benchmark. TERMINAL-BENCH 4.0 · AGENTIC TASKS 10.3% Claude Sonnet 5 66.4% Claude Opus 5.5 70.6% Claude Sonnet 5.5 genztech.blog
Fig 1 Sonnet 5.5's Terminal-Bench 4.0 score isn't just an improvement over Sonnet 5, it edges out Anthropic's own more expensive Opus 5.5 on the same test.

What actually changed between Sonnet 5 and Sonnet 5.5?

Anthropic is pitching this release almost entirely around agentic reliability rather than raw intelligence. Beyond the Terminal-Bench jump, Sonnet 5.5 scores 46.2% on FrontierCode 1.1 at max effort, 55.5% on CursorBench 4.0, and 1844 on GDPval-AA v2.1, a hair behind Opus 5.5's 1846 on the same scale. Anthropic also says outputs generate more than 30% faster than Sonnet 5, and that this is the first Sonnet-tier model to ship with cyber safeguards comparable to what Opus 5.5 gets, a tier of protection Anthropic previously reserved for its flagship. One detail Anthropic highlighted that's easy to verify yourself: Sonnet 5.5 is the first Sonnet model that can beat Pokemon Red working only from screenshots, a long-running informal test of whether a model can plan and adapt over hundreds of steps without a human correcting its course.

RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price

Why does a benchmark most people have never heard of matter here?

SWE-bench Verified, the benchmark that made Sonnet 5's June launch a story, measures whether a model can patch a single, well-defined GitHub issue correctly. Terminal-Bench 4.0 measures something closer to how these models actually get used inside tools like Claude Code, Cursor and GitHub Copilot Workspace: long sequences of terminal commands, file edits and course corrections strung together with nobody checking each individual step. A model can be excellent at patching one clean PR and still fall apart forty steps into an unsupervised agent run, which is roughly what Sonnet 5's 10.3% score suggests happened. Anthropic didn't publish any SWE-bench Verified number for Sonnet 5.5 at all, the same choice it made for Opus 5.5 three weeks ago. That's a deliberate pivot, not an oversight, and it lines up with a separate fact worth knowing: vals.ai, the independent lab that has been the field's most trusted SWE-bench evaluator, archived its SWE-bench Verified leaderboard on September 5 and isn't expected to score new models on it going forward.

Is it actually cheaper, or just faster?

Worth being precise here, because "30% cheaper" is doing a lot of marketing work. The rate card is identical to Sonnet 5's: $2 per million input tokens, $10 per million output, $0.20 for cached reads, $2.50 for cache writes. Nothing on that list moved. What Anthropic is actually claiming is a total-cost-per-finished-task number: because Sonnet 5.5 is over 30% faster and needs fewer output tokens to reach a working answer, the same job costs less to complete even though every individual token is billed at the old rate. It's the same shape of claim Anthropic made when Opus 5.5 launched, where a "40% less to run" figure folded in a 30% speed gain rather than describing an actual rate cut. Neither claim is false, but neither is a discount either. If your usage pattern doesn't change, your bill per token doesn't change.

SpecClaude Sonnet 5Claude Sonnet 5.5Claude Opus 5.5
List price (in / out per 1M)$2 / $10$2 / $10$4 / $20
Terminal-Bench 4.010.3%70.6%66.4%
SWE-bench Verified85.2% (independent, vals.ai)Not publishedNot published
ReleasedJune 30, 2026September 28, 2026September 22, 2026
Best forDefault free/Pro chat modelAgentic coding at Sonnet-tier pricingHardest agentic and computer-use work

What does this mean for the market?

Anthropic isn't public, so there's no ticker to watch directly, but the ripple effects land on its cloud partners and its rivals. AWS, Google Cloud and Microsoft Azure all host Sonnet 5.5 at launch, and every one of them earns from Claude usage regardless of whether Sonnet or Opus wins a given workload, so a cheaper-feeling mid-tier model that gets adopted more widely is a net positive for all three. For enterprise buyers, the number that matters isn't the benchmark score, it's what happens when you multiply a modest per-task efficiency gain by an agent fleet running thousands of calls a day; even Anthropic's own examples lean into this, with Zendesk reporting support tickets processed noticeably faster and Unity citing a near-complete pass rate on its internal multi-step benchmark. The competitive pressure is real and immediate too. TechCrunch's coverage of the launch notes OpenAI shipped updated versions of its Sol and Luna models the week before, and Meta has been pushing its own agents into smart-glasses features. Anthropic did the same thing in June when Sonnet 5 reset the price-to-performance curve for coding models; the signal for anyone watching this space is that the mid-tier is where the real competition is happening now, not at the top of each lab's lineup.

RelatedClaude Opus 5.5 Ships With Agent Benchmarks, Not SWE-bench

Who's actually using it already?

Anthropic's launch materials name a wider set of early customers than usual. Epic Games' COO Daniel Vogel said the model cleared the quality bar his team expects from a higher tier. Zendesk's Abhinay Kathuria pointed to a roughly 20% speedup in ticket processing. SpaceXAI's Sualeh Asif called out the CursorBench 4.0 number specifically as frontier-level for a mid-tier model, and Unity's Sam Zhang said it completed nine out of ten tasks in the studio's internal multi-step benchmark. Slack, Atlassian, Box, Lovable and CodeRabbit are also listed among launch partners, which reads less like a curated highlight reel and more like Anthropic wanting to show this model is already load-bearing inside real products, not just a benchmark exercise.

What to watch · coming weeks
  • Whether an independent SWE-bench Verified score ever appears. With vals.ai's board archived since September 5, there's no obvious path to a neutral number for this model on the metric that made Sonnet 5's launch credible.
  • How OpenAI and Google respond in the mid-tier. Sonnet 5 forced a price-to-performance reset in June; if Sonnet 5.5's Terminal-Bench result holds up under independent testing, expect a competing mid-tier release within weeks, not months.
  • The new Haiku model Anthropic says is coming this quarter. TechCrunch's report mentions it in passing; a cheap, fast Haiku built on the same Terminal-Bench-focused approach would complete the Claude 5.5 lineup and matters more to high-volume agent deployments than either Sonnet or Opus alone.

Our take

The Terminal-Bench number is the real story here, and it's a genuinely large jump that says more about how Sonnet 5.5 was built than any marketing line does: Anthropic optimized specifically for long, unsupervised agent runs, and it shows. The "30% cheaper" framing is honest but easy to misread as a price cut it isn't; nothing on the rate card moved, and the savings only materialize if your workload actually benefits from the speed. What's missing, and what's been missing since Opus 5.5 three weeks ago, is any independent SWE-bench number to check Anthropic's benchmark suite against, since the neutral evaluator that used to do that job stopped in September. Until something fills that gap, Sonnet 5.5's numbers deserve the same read as any vendor's: plausible, consistent with a real architectural improvement, and unverified.

Primary sources

Original analysis by GenZTech Team.