head-to-head
| Metric | GPT-5.6 Sol | Gemini 3.7 Flash |
|---|---|---|
| SWE-bench Verified | 96.2% | 80.8% |
| SWE-bench Pro | — | — |
| Terminal-Bench | 88.8% | — |
| Input $ / 1M | $5 | $0.75 |
| Output $ / 1M | $30 | $3.75 |
| Context | — | — |
| Open weights | No | No |
| Access | API · Codex (public since Jul 9 2026) | API (gemini-3.7-flash) · AI Studio · Google Antigravity · Android Studio · Gemini app Spark (AI Pro/Ultra) · Gemini Enterprise |
| Maker | OpenAI | Google DeepMind |
what do the benchmarks actually say?
On SWE-bench Verified — real, human-validated GitHub issues resolved end-to-end — GPT-5.6 Sol posts 96.2% against 80.8% for Gemini 3.7 Flash, a 15.4-point gap. Verified is the closest public proxy for "can it fix a real bug in a real repo without help", which is why it anchors our ranking.
A few points either way is real but not decisive: within that band, the agent scaffolding around the model — how it retrieves files, runs tests, and retries — often matters as much as the base model. Treat the gap as a lean, not a verdict.
which is cheaper to run?
Gemini 3.7 Flash is the cheaper model: $0.75 per 1M input tokens ($3.75 output) versus $5 ($30 output) for GPT-5.6 Sol — roughly 6.7× less on input. Coding workloads are output-heavy — agents write diffs, tests and retries — so weight the output rate more than the input rate when you estimate a monthly bill.
when to pick each
The strongest model on long tasks: 98% on the 1-to-4-hour tier, ahead of Claude Opus 5, and the second-highest overall score.
Google's volume tier at half the price of 3.6 Flash, and the first 3.7-series model with an independent coding score. A real gain over 3.6 Flash, unlike the last Flash bump.
how were these scores verified?
We only print a number once it's confirmed against a primary source or an independent evaluation, and each row on our leaderboard records which kind it is:
- GPT-5.6 Sol: Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board. Correction, Jul 25, 2026: this row read "the top score on the board" from Jul 17 until Jul 25, when vals.ai evaluated Claude Opus 5 at 97.00% ±0.76 and took the #1 slot. The 0.8-point gap is ~0.7 sigma and not significant, so the two are a statistical tie, and Sol still leads on the longest tasks (98% vs 90% on the 1-to-4-hour tier). Verified Jul 17, 2026; it had been unranked since Jun 26 because OpenAI published no SWE-bench number of its own, and it still has not. Read the #1 with care: the 1.2-point lead over Claude Fable 5 (95.00% ±0.98) is inside the combined margin of error (~0.9 sigma, not significant), so the two are a statistical tie and we rank Sol first only because it scored higher. Where it does separate is task length — 98% on 1-4 hour tasks vs 93% for Fable 5. OpenAI's own Terminal-Bench 2.1 claim is 88.8% (Sol) / 91.9% (Sol Ultra). No SWE-bench Pro score published. Pricing $5/$30 per 1M.
- Gemini 3.7 Flash: Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 80.80% ±1.76. Released Aug 13, 2026, three weeks after Gemini 3.6 Flash. Google published NO SWE-bench Verified score of its own, exactly as with 3.6 Flash, so this rank rests entirely on the independent run. Its launch claims are all vendor-reported on Google scaffolding: DeepSWE v1.1 65.3% (up from 49.0%), FrontierCode 1.1 Main 43.6% (up from 34.4%), WebDev Arena Elo 1588 (up from 1538), GDP.pdf 34.0% (up from 22.0%), AutomationBench 30.4% (up from 17.0%). Read those with the 3.6 Flash precedent in mind: Google claimed a 12-point DeepSWE jump for that model too, and vals.ai then measured it at 79.60% ±1.80 against 78.8% for 3.5 Flash, a statistical tie. Price: introductory $0.75/$3.75 per 1M through Dec 31, 2026, reverting to $1.50/$7.50 on Jan 1, 2027, i.e. back to 3.6 Flash list. Specs: 1,048,576 input / 65,536 output tokens; thinking levels low/medium/high only (minimal now errors); computer use still preview. See /p/gemini-3-7-flash-half-price-3-5-pro-missing/. **Ranked 2026-08-17**: vals.ai evaluated it and we now rank on that independent figure, since Google still publishes no SWE-bench Verified score of its own. It entered our verifying queue on Aug 13 and left it four days later. Worth recording that the 3.6 Flash precedent did NOT repeat: that model turned a claimed 12-point DeepSWE jump into a statistical tie, whereas 3.7 Flash gains a real 1.2 points over 3.6 Flash (79.60%) on the same harness. The gain is still inside the combined margin of error, so read 3.7 Flash, 3.6 Flash and Claude Sonnet 5 as a three-way tie rather than a ranking.
Full reviewsGPT-5.6 Sol, decodedGemini 3.7 Flash, decoded
Ranked on our AI Coding Leaderboard, updated 2026-08-17. Scores are confirmed against primary sources; prices are per 1M input tokens and can change.
- OpenAIvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board. Correction, Jul 25, 2026: this row read "the top score on the board" from Jul 17 until Jul 25, when vals.ai evaluated Claude Opus 5 at 97.00% ±0.76 and took the #1 slot. The 0.8-point gap is ~0.7 sigma and not significant, so the two are a statistical tie, and Sol still leads on the longest tasks (98% vs 90% on the 1-to-4-hour tier). Verified Jul 17, 2026; it had been unranked since Jun 26 because OpenAI published no SWE-bench number of its own, and it still has not. Read the #1 with care: the 1.2-point lead over Claude Fable 5 (95.00% ±0.98) is inside the combined margin of error (~0.9 sigma, not significant), so the two are a statistical tie and we rank Sol first only because it scored higher. Where it does separate is task length — 98% on 1-4 hour tasks vs 93% for Fable 5. OpenAI's own Terminal-Bench 2.1 claim is 88.8% (Sol) / 91.9% (Sol Ultra). No SWE-bench Pro score published. Pricing $5/$30 per 1M.
- Google DeepMindvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 80.80% ±1.76. Released Aug 13, 2026, three weeks after Gemini 3.6 Flash. Google published NO SWE-bench Verified score of its own, exactly as with 3.6 Flash, so this rank rests entirely on the independent run. Its launch claims are all vendor-reported on Google scaffolding: DeepSWE v1.1 65.3% (up from 49.0%), FrontierCode 1.1 Main 43.6% (up from 34.4%), WebDev Arena Elo 1588 (up from 1538), GDP.pdf 34.0% (up from 22.0%), AutomationBench 30.4% (up from 17.0%). Read those with the 3.6 Flash precedent in mind: Google claimed a 12-point DeepSWE jump for that model too, and vals.ai then measured it at 79.60% ±1.80 against 78.8% for 3.5 Flash, a statistical tie. Price: introductory $0.75/$3.75 per 1M through Dec 31, 2026, reverting to $1.50/$7.50 on Jan 1, 2027, i.e. back to 3.6 Flash list. Specs: 1,048,576 input / 65,536 output tokens; thinking levels low/medium/high only (minimal now errors); computer use still preview. See /p/gemini-3-7-flash-half-price-3-5-pro-missing/. **Ranked 2026-08-17**: vals.ai evaluated it and we now rank on that independent figure, since Google still publishes no SWE-bench Verified score of its own. It entered our verifying queue on Aug 13 and left it four days later. Worth recording that the 3.6 Flash precedent did NOT repeat: that model turned a claimed 12-point DeepSWE jump into a statistical tie, whereas 3.7 Flash gains a real 1.2 points over 3.6 Flash (79.60%) on the same harness. The gain is still inside the combined margin of error, so read 3.7 Flash, 3.6 Flash and Claude Sonnet 5 as a three-way tie rather than a ranking.
- BenchmarkSWE-bench — the real-GitHub-issue benchmark