specs at a glance

Leaderboard rank#6 of 22
SWE-bench Verified88.8%
SWE-bench Pro
Terminal-Bench82.7 (TB2.1, vendor)
Input price / 1M$0.14
Output price / 1M$0.28
Context window1M
Open weightsYes
AccessOpen weights (MIT) · API (deepseek-v4-flash) · self-host
MakerDeepSeek

how good is DeepSeek V4 Flash at coding?

DeepSeek V4 Flash sits at #6 of 22 ranked models, posting 88.8% on SWE-bench Verified — 8.2 points behind #1 Claude Opus 5. Terminal-Bench (agentic terminal work): 82.7 (TB2.1, vendor).

Score provenance: Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the checkpoint DeepSeek own change log describes. Ranks above Claude Opus 4.8 (88.6%) and below GPT-5.6 Luna (93.0%). Vendor-reported agent numbers for the same build (DeepSeek change log, Jul 31 2026): Terminal-Bench 2.1 82.7, Cybergym 76.7, Toolathlon verified 70.3, DeepSWE 54.4, NL2Repo 54.2, Agent Last Exam 25.2, produced with DeepSeek own unreleased "DeepSeek Harness minimal mode" and not independently reproduced. The V4-Flash-Preview model card (pre-0731 weights) separately listed SWE-bench Verified 79.0 and Terminal-Bench 2.0 56.9; those are not carried over. Same architecture and size as the preview (284B total / 13B active, FP4+FP8), re-post-trained only. Not to be confused with DeepSeek V4 Pro (ranked, 80.6%) or the plain DeepSeek V4 (77.4%). See /p/deepseek-v4-flash-claims-82-7-terminal-bench/.

what does DeepSeek V4 Flash cost?

$0.14 per 1M input tokens and $0.28 per 1M output — the cheapest model we track. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.

where can you use it?

Available via Open weights (MIT) · API (deepseek-v4-flash) · self-host. Because it ships open weights, you can also self-host it on your own hardware or any inference provider — with the version pinned so the model can't change under you.

head-to-head

Full storyDeepSeek V4 Flash claims 82.7 on Terminal-Bench 2.1

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-08-06.