the ranking
| # | Model | $/1M input | Best for |
|---|---|---|---|
| 1 | Laguna S 2.1 openPoolside | $0.10/1M | Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware. |
| 2 | DeepSeek V4 Flash openDeepSeek | $0.14/1M | The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output. |
| 3 | GPT-5.6 LunaOpenAI | $0.20/1M | The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token. |
| 4 | Gemini 3.5 Flash-LiteGoogle DeepMind | $0.30/1M | The cheapest ranked model on the board: 75.0% independently measured at $0.30 per 1M input, roughly a fifth the price of the Flash tier above it. |
| 5 | LongCat-2.0 openMeituan | $0.30/1M | A 1.6T Mixture-of-Experts agentic coder trained entirely on domestic Chinese chips; led OpenRouter usage in stealth as "Owl Alpha". |
| 6 | DeepSeek V4 Pro openDeepSeek | $0.435/1M | The cheapest frontier-class coder — top open-weights score at ~11× less than Opus. Best pick when cost or self-hosting rules. |
the top picks, decoded
1. Laguna S 2.1 — $0.10/1M $/1M input
Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware. Vendor-reported (Poolside, Jul 21 2026): Terminal-Bench 2.1 70.2% with thinking enabled (60.4% without), SWE-bench Pro 59.4%, SWE-bench Multilingual 78.5%, DeepSWE v1.1 40.4%. Poolside published no SWE-bench Verified score, so it stays unranked pending an independent eval. vals.ai has evaluated Laguna M.1 (57.6%) and Laguna XS.2 (55.2%) but not S 2.1, and a near-miss version is not a match. Price: OpenRouter $0.10/$0.20 per 1M. Re-checked 2026-08-16: vals.ai has not evaluated it.
2. DeepSeek V4 Flash — $0.14/1M $/1M input
The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output. Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the checkpoint DeepSeek own change log describes. Ranks above Claude Opus 4.8 (88.6%) and below GPT-5.6 Luna (93.0%). Vendor-reported agent numbers for the same build (DeepSeek change log, Jul 31 2026): Terminal-Bench 2.1 82.7, Cybergym 76.7, Toolathlon verified 70.3, DeepSWE 54.4, NL2Repo 54.2, Agent Last Exam 25.2, produced with DeepSeek own unreleased "DeepSeek Harness minimal mode" and not independently reproduced. The V4-Flash-Preview model card (pre-0731 weights) separately listed SWE-bench Verified 79.0 and Terminal-Bench 2.0 56.9; those are not carried over. Same architecture and size as the preview (284B total / 13B active, FP4+FP8), re-post-trained only. Not to be confused with DeepSeek V4 Pro (ranked, 80.6%) or the plain DeepSeek V4 (77.4%). See /p/deepseek-v4-flash-claims-82-7-terminal-bench/.
3. GPT-5.6 Luna — $0.20/1M $/1M input
The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token. Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Pricing added Aug 1, 2026: OpenAI cut Luna 80% on Jul 30, 2026, from $1/$6 to $0.20/$1.20 per 1M, making it the cheapest ranked model on this board per token. Vendor-announced (OpenAI, via its own post and consistent reporting from BleepingComputer and VentureBeat); openai.com/api/pricing returns 403 to our fetcher, so this is a vendor figure rather than one we read off the pricing page ourselves. It had been left blank since Jul 21 for exactly that reason.
how we rank
We rank by SWE-bench Verified (500 real, human-validated GitHub issues resolved end-to-end), tiebroken by the harder SWE-bench Pro. A score is only printed once confirmed against an independent evaluation or the maker's primary source — and every row states which kind it is. Where both exist, we print both: one as the ranked score, the other in that row's note. We would rather show you the gap than ask you to trust our pick. Our independent reference is vals.ai, which runs every model itself through the same minimal bash-only harness (mini-swe-agent), so the models are compared on equal footing. That matters more than it sounds: SWE-bench scores a model and its scaffolding together, and vendors report using their own tuned scaffolds. Against vals.ai's neutral harness, the vendor claims on this board run 2.6 to 11.6 points optimistic. So rows marked vendor-reported are best-case numbers and are not strictly comparable to the independent ones — where we know the independent figure, we print it in the row's note. That is a deliberate choice and you should know we made it: on four rows (Claude Sonnet 5, MiniMax M3, Qwen3.7 Max, Kimi K2.6) an independent score for that exact model exists and is lower, and we still rank on the maker's published figure because it is the number that model is sold and quoted on. We disclose the independent one beside it instead of quietly restating the board on a single evaluator's harness choice. The honest consequence: positions that straddle the two regimes are approximate. Qwen3.7 Max is the sharpest case — it sits at #16 on Alibaba's 80.4%, and on the neutral harness its 68.8% would put it far down the table. Note that llm-stats, which we previously miscredited as an independent tracker, labels its own SWE-bench Verified table "Verified: 0 / Self-reported: 104"; it aggregates vendor claims. Models still being checked are marked “verifying” and shown without a number rather than estimated. Prices are per 1M input tokens on the standard API tier and can change — always confirm current pricing with the provider.
Want the raw numbers? The full dataset is public: JSON · CSV.
From our full AI Coding Leaderboard (2026-08-13). We only rank scores confirmed against primary sources.