specs at a glance

Leaderboard rank#7 of 28
SWE-bench Verified93.4%
SWE-bench Pro
Terminal-Bench88.3%
Input price / 1M$3
Output price / 1M$15
Context window1M
Open weightsYes
AccessOpen weights · API · app
MakerMoonshot AI

how good is Kimi K3 at coding?

Kimi K3 sits at #7 of 28 ranked models, posting 93.4% on SWE-bench Verified — 3.6 points behind #1 Claude Opus 5. Terminal-Bench (agentic terminal work): 88.3%.

Score provenance: Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 93.40% ±1.11. Verified Jul 18, 2026 — vals.ai had not run K3 at our Jul 17 check and now ranks it third overall (above GPT-5.6 Luna at 93.0% and Claude Opus 4.8 at 88.6%), making it the highest-scoring open-weight coder we track. Moonshot published no SWE-bench Verified score of its own, only Terminal-Bench 2.1 88.3% on its KimiCode harness, so we rank on the independent number per our standing rule. Corroborated by a second independent evaluator: Artificial Analysis, which runs its own tests, has K3 4th on GDPval-AA v2 (Elo 1684), 2nd on AA-Briefcase (Elo 1545) and 4th of 187 on its Intelligence Index v4.1 (57.1). It also classifies K3 as proprietary outright. AA measured Terminal-Bench 2.1 at 85.02% against Moonshot's own claim of 88.3%, a 3.3-point vendor overstatement, which is why we print the independent figure. Blind developer voting on Arena WebDev ranks it 1st of 99. Price confirmed on Moonshot's own pricing page: $3/$15 per 1M, cache-hit input $0.30 (platform.kimi.ai, checked Jul 19, 2026). **open: false, corrected Jul 19, 2026** — we had this flagged as open weights, which put it top of our "best open-source coding model" page as a model nobody can download. Checked Jul 19: no K3 repo under huggingface.co/moonshotai (18 repos, none K3) and github.com/MoonshotAI/Kimi-K3 returns 404. Moonshot's own blog still says "The full model weights will be released by July 27, 2026." Until they actually ship, open is a roadmap item, not a fact, and this row is an API model. Flip back to open: true when the weights land. Re-checked Jul 27, 2026 (Moonshot's own deadline day): huggingface.co/moonshotai still lists no K3 repo, newest is Kimi-K2.7-Code. Score unchanged at 93.40% on vals.ai. **open: true, restored Jul 27, 2026 13:31 UTC**: the weights actually landed at huggingface.co/moonshotai/Kimi-K3 later the same day, 96 safetensors shards, 1,560,998,983,759 bytes (about 1,454 GiB), under a bespoke "Kimi K3 License" that is permissive but not Apache 2.0 (separate agreement required above $20M MaaS revenue; attribution required above 100M MAU). Verified against the Hugging Face API and config.json, not a press report. One caveat that belongs on the record: the release is 4-bit MXFP4 only for the routed experts (quantization-aware trained from the SFT stage, format mxfp4-pack-quantized, group size 32), with attention, shared experts, LM head and vision tower left at bf16, and there is no base checkpoint, so this is open for inference rather than open for training. Score unchanged at 93.40%; nothing was re-ranked by this flag. See /p/kimi-k3-weights-landed-4-bit-only/.

what does Kimi K3 cost?

$3 per 1M input tokens and $15 per 1M output — #21 cheapest of the 26 models we track. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.

where can you use it?

Available via Open weights · API · app. Because it ships open weights, you can also self-host it on your own hardware or any inference provider — with the version pinned so the model can't change under you.

Full storyKimi K3 Shipped Early. The Real Barrier Isn't the Repo.

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-08-20.