specs at a glance
| Leaderboard rank | #4 of 15 |
|---|---|
| SWE-bench Verified | 93.0% |
| SWE-bench Pro | — |
| Terminal-Bench | — |
| Input price / 1M | — |
| Output price / 1M | — |
| Context window | — |
| Open weights | No |
| Access | API |
| Maker | OpenAI |
how good is GPT-5.6 Luna at coding?
GPT-5.6 Luna sits at #4 of 15 ranked models, posting 93.0% on SWE-bench Verified — 3.2 points behind #1 GPT-5.6 Sol.
Score provenance: Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Per-token list pricing not confirmed against OpenAI's own pricing page, so inPrice/outPrice stay blank rather than estimated.
where can you use it?
Available via API. As a proprietary model, you're on the maker's infrastructure and release schedule.
head-to-head
Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-07-21.