DeepSeek pushed V4-Flash out of preview and into public beta a few hours ago, and attached to the release a run of agent benchmark scores led by 82.7 on Terminal-Bench 2.1. That would put a model priced at $0.28 per million output tokens inside seven points of the strongest coding agents on the market. The asterisk is large: DeepSeek measured it with something it calls the DeepSeek Harness in minimal mode, and that harness has not shipped yet, so nobody outside the company can reproduce the number today.
The build is labelled DeepSeek-V4-Flash-0731, and the changelog is unusually specific about what did not change. Same architecture, same size as the preview. The company says it was only re-post-trained. Whatever moved these scores, it was not more parameters and not a longer pretraining run.
RelatedLaguna S 2.1: 8B Active Params, 70% Terminal-Bench
- The API model name is unchanged. Point at
deepseek-v4-flashand you get the new weights, no migration step. - Published agent scores: Terminal-Bench 2.1 at 82.7, Cybergym 76.7, Toolathlon verified 70.3, NL2Repo 54.2, DeepSWE 54.4, and a much harsher 25.2 on Agent Last Exam.
- V4-Flash now speaks the Responses API format natively and is specifically adapted for Codex, which is a distribution move as much as a model release.
- V4-Pro and the app and web models are untouched. DeepSeek says the official V4-Pro release follows soon.
What actually changed in the 0731 build?
Post-training, and apparently a lot of it. The preview card on Hugging Face lists V4-Flash as a 284B-parameter mixture of experts with 13B active per token, a 1M-token context window, and mixed FP4 and FP8 precision under an MIT license. It scored 56.9 on Terminal-Bench 2.0 and 79.0 on SWE-bench Verified. Today's build keeps every one of those architectural facts and reports 82.7 on Terminal-Bench 2.1.
The two Terminal-Bench versions are not the same test, so that is not a clean 26-point delta. Version 2.1 is described as a verified refresh of 2.0 with environment and instruction fixes, which tends to lift scores across the board by removing broken tasks. Even discounting for that, a jump of this size out of re-post-training alone is the interesting part of the announcement. DeepSeek is claiming the agent gap was a training-recipe problem, not a capacity problem.
The company also says the new Flash beats V4-Pro-Preview, its own larger and more expensive tier, on these agent tests. That inverts DeepSeek's product ladder, at least until the official Pro build lands.
How does 82.7 compare to what it is chasing?
Artificial Analysis runs its own Terminal-Bench v2.1 evaluation using the Terminus 2 agent harness inside an e2b sandbox, pass@1 averaged over three repeats per task. On that board the leaders are GPT-5.6 Sol at 89.5%, Claude Opus 5 at 89.1%, and GPT-5.6 Terra at 88.0%. DeepSeek's self-reported 82.7 would sit just under that group.
Why the harness matters more than the number
Terminal-Bench does not score a model in isolation. It scores a model plus the scaffolding that feeds it the terminal, retries its mistakes, and decides when it is done. Change the harness and the same weights post a different number. This is the single most common way benchmark claims mislead, and it is why our AI coding leaderboard tracks vendor and independent figures separately: across the models where both exist, vendor-run numbers have come in 2.6 to 11.6 points above a neutral bash-only harness.
DeepSeek's footnote is candid about the setup, which helps. It ran the code-agent tasks through DeepSeek Harness minimal mode at max effort with top-p 0.95 and temperature 1.0, and notes the harness is to be released soon. Two of the nine reported results, DSBench-FullStack at 68.7 and DSBench-Hard at 59.6, come from internal test sets that nobody else can score against at all.
None of that makes 82.7 wrong. It makes it unverified. The honest position until Artificial Analysis or vals.ai runs the 0731 weights on a neutral harness is that DeepSeek has published a strong claim with its methodology disclosed and its tooling withheld.
What does Codex support actually buy DeepSeek?
More than the benchmark line, probably. Native Responses API support plus explicit Codex adaptation means a developer already wired into OpenAI's agent tooling can swap the model without rewriting the integration. DeepSeek has done this before with the Anthropic-compatible endpoint it shipped alongside V4 in April. The pattern is consistent: match whatever interface the incumbent tools already speak, then compete on price.
RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price
| V4-Flash | V4-Pro | |
|---|---|---|
| Input, cache miss | $0.14 / 1M | $0.435 / 1M |
| Input, cache hit | $0.0028 / 1M | $0.003625 / 1M |
| Output | $0.28 / 1M | $0.87 / 1M |
| Context | 1M tokens | 1M tokens |
| Max output | 384K tokens | 384K tokens |
| Status today | 0731 build, public beta | Preview, official build pending |
Read the output column next to the benchmark claim and you get the actual pitch. Flash costs roughly a third of Pro per output token and now claims to beat Pro-Preview on agent work. Agent workloads are output-heavy and loop for thousands of tokens per task, so that column is where the bill is decided.
One piece of fine print worth pricing in: DeepSeek's docs say the API will soon move to peak and off-peak rates, with peak hours costing double. The windows given are 9:00 to 12:00 and 14:00 to 18:00 Beijing time. A US team running overnight batch jobs lands in off-peak by accident. A team in Singapore or Bangalore does not.
What it means for the market
The exposure here is not to DeepSeek, which is private and does not sell equity to you. It is to the price floor for agentic inference. Every credible sub-dollar model that posts near-frontier agent scores compresses what OpenAI, Anthropic and Google can charge for the tier below their flagships, and that tier is where most production agent traffic actually runs. Watch Microsoft and Alphabet commentary on inference gross margin rather than any single model launch, and watch whether the Codex and Claude Code ecosystems start surfacing third-party model backends as a first-class option. The signal for investors is margin compression in mid-tier inference, not a change in who holds the top of the leaderboard.
- Dec 2025V3.2 ships deepseek-chat and deepseek-reasoner upgraded
- Apr 24, 2026V4 preview V4-Pro and V4-Flash land, OpenAI and Anthropic interfaces
- Jul 24, 2026Legacy names retired deepseek-chat and deepseek-reasoner switched off
- Jul 31, 2026V4-Flash-0731 Public beta, agent scores published, Codex support
- SoonDeepSeek Harness release Needed before anyone can reproduce 82.7
- SoonV4-Pro official build Company says it follows this release
Our take
The claim we would most like checked is not the Terminal-Bench score. It is the statement that architecture and size did not move. If a 284B model with 13B active parameters really picked up this much agent capability from re-post-training alone, that is a more useful finding for everyone building on open weights than any single leaderboard row, because post-training recipes travel and pretraining runs do not.
We are not adding a ranked score for V4-Flash to our leaderboard today. DeepSeek published no SWE-bench Verified figure for the 0731 build, and the preview's 79.0 belongs to different weights. It goes on the board as a verifying row with the Terminal-Bench claim recorded and no rank, which is the same treatment every unconfirmed number gets there.
- The harness release. Until DeepSeek Harness minimal mode is public, 82.7 cannot be reproduced by anyone. If it stays unreleased for long, treat the number accordingly.
- A neutral rerun. Artificial Analysis and vals.ai both run their own scaffolds. Whichever gets to the 0731 weights first sets the real figure.
- The V4-Pro build. DeepSeek says it is next. If Pro does not clearly beat this Flash on agent work, the tiering stops making sense.
- Peak pricing. The 2x peak window is announced but not dated. It changes the cost case for anyone in an Asian timezone.
- OfficialDeepSeek API change log — the 2026-07-31 V4-Flash entry and its benchmark footnotes
- OfficialDeepSeek models and pricing — per-token rates, context and the peak-hour note
- ReferenceDeepSeek-V4-Flash model card — 284B/13B MoE, MIT license, preview benchmark table
- BenchmarkArtificial Analysis, Terminal-Bench v2.1 — independent scores on the Terminus 2 harness
- DataGENZ TECH AI coding leaderboard — our vendor-versus-independent scoring rules and the V4 rows
Original analysis by GenZTech, built from DeepSeek's own release notes and pricing documentation, cross-checked against independent Terminal-Bench v2.1 results.
