specs at a glance

Leaderboard statusVerifying — no confirmed SWE-bench Verified score
SWE-bench Verified
SWE-bench Pro
Terminal-Bench66.4 (TB4, vendor)
Input price / 1M$4
Output price / 1M$20
Context window
Open weightsNo
AccessAPI (claude-opus-5-5) · Claude Platform · AWS · Google Cloud · Microsoft Azure
MakerAnthropic

how good is Claude Opus 5.5 at coding?

Claude Opus 5.5 is on our board but deliberately unranked. We rank by SWE-bench Verified, Anthropic has not published that number, and no independent evaluation has run it, so there is nothing here we could stand behind. What the maker does publish is 66.4 (TB4, vendor) on Terminal-Bench, measured on its own scaffold and not comparable like-for-like with the ranked rows. It moves into the ranking the moment a confirmed score exists. Terminal-Bench (agentic terminal work): 66.4 (TB4, vendor).

Score provenance: Vendor-reported (Anthropic, Sept 22 2026): Terminal-Bench 4.0 66.4%, FrontierCode v1.1 54.4%, CursorBench 4.0 57.8%, GDPval-AA v2.1 1846 Elo, AutomationBench 40.0%, OSWorld 2.0 81.8%, Chartography 89.0%. Anthropic published NO SWE-bench Verified score for this model, so there is nothing to rank on our tracked metric and it enters the verifying queue unranked.

Pricing precision: list price is $4/$20 per 1M against Opus 5's $5/$25, i.e. 20% lower per token; Anthropic's own claim of '40% less to run than Opus 5' is a total-cost claim that also folds in a stated 30% faster output, not a per-token discount, so the two figures are not in conflict but are not the same measure either. Queue caveat: vals.ai archived its SWE-bench Verified board on 5 September 2026, so this row is unlikely to ever receive an independent score from that source.

See /p/claude-opus-5-5-terminal-bench-agentic-benchmarks/. Re-checked 2026-09-23: vals.ai has not evaluated it; its SWE-bench Verified board is archived (88 systems, frozen since 2026-09-05) and the nearest names on it are Claude Opus 5 and Claude Opus 4.8, neither of which is this model.

what does Claude Opus 5.5 cost?

$4 per 1M input tokens and $20 per 1M output — as listed by the maker. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.

where can you use it?

Available via API (claude-opus-5-5) · Claude Platform · AWS · Google Cloud · Microsoft Azure. As a proprietary model, you're on the maker's infrastructure and release schedule.

Full storyClaude Opus 5.5 Ships With Agent Benchmarks, Not SWE-bench

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-09-03.