Google shipped Gemini 3.7 Flash at 10 AM Pacific this morning, three weeks to the day after 3.6 Flash, and cut the price in half on the way out the door: $0.75 per million input tokens and $3.75 per million output, held there through December 31, 2026. The pitch is coding and agent work, and Google's own numbers move in double digits. DeepSWE v1.1 goes from 49.0% to 65.3%. FrontierCode 1.1 Main goes from 34.4% to 43.6%.
What Google did not ship, for the third launch running, is Gemini 3.5 Pro.
RelatedDeepSeek V4 Pro hits GA, but Flash still scores higher
- The price is the story. $0.75 in and $3.75 out per million tokens is exactly half of what 3.6 Flash cost at launch, and the intro rate runs 140 days before doubling to $1.50 and $7.50 on January 1, 2027.
- Every published benchmark jumped. DeepSWE v1.1 49.0% to 65.3%, FrontierCode 1.1 Main 34.4% to 43.6%, WebDev Arena Elo 1538 to 1588, GDP.pdf 22.0% to 34.0%, AutomationBench 17.0% to 30.4%.
- Google published no SWE-bench Verified score, same as with 3.6 Flash. Our AI coding leaderboard holds 3.7 Flash unranked until a neutral harness runs it.
- The Flash line is now two version numbers ahead of a Pro flagship that has never been released to anyone outside a partner test.
What actually changed from 3.6 Flash?
The model code is gemini-3.7-flash. It takes a 1,048,576 token input window and returns up to 65,536 tokens, accepts text, images, video, audio and PDFs, and keeps the tool surface Google has been building out all year: caching, code execution, file search, function calling, structured outputs, Search grounding, Maps grounding and computer use, still marked preview.
One change will break existing code. Thinking levels are now low, medium and high only. Passing minimal returns an error rather than falling back, so anything in your codebase that pinned the cheapest reasoning tier on a Flash model needs an edit before you swap the model string.
Google's framing for the release is agent reliability rather than raw capability. The company says the model shows "a more disciplined execution" that means "less manual oversight and fewer retries across engineering workflows," which is a claim about wasted tokens as much as about correctness. The launch post also flags updated safeguards in CBRN and cyber offense domains, the same category of guardrail Google leaned on when it announced Gemini 3.5 Flash Cyber in July and then declined to give anyone access to it.
How much of that benchmark jump should you believe?
Here is the part worth slowing down for, because we have the receipts from last time. When Google launched 3.6 Flash on July 21, it claimed DeepSWE had gone from 37% to 49%, a twelve point jump on the same benchmark it is now quoting again. Two days later vals.ai ran 3.6 Flash on a neutral mini-swe-agent bash-only harness and measured 79.60% ±1.80 on SWE-bench Verified, against 78.8% for the older 3.5 Flash. That is a 0.8 point gap, comfortably inside the combined error bars. A twelve point vendor jump landed as a statistical tie on the neutral harness.
That is not an accusation that Google is inventing numbers. It is the well documented gap between benchmarking a model and benchmarking a model plus the scaffolding its maker wrote for it. Across our leaderboard, vendor-reported coding scores run 2.6 to 11.6 points above the independent figure. DeepSWE is also Google's own benchmark, which makes it a fine measure of progress against Google's last model and a poor one for comparing against Anthropic's or OpenAI's.
So the honest read on 3.7 Flash today: the price cut is a fact you can act on immediately, and the coding gains are a claim you should wait roughly two days to price in.
| Gemini 3.5 Flash | Gemini 3.6 Flash | Gemini 3.7 Flash | |
|---|---|---|---|
| Launched | May 19, 2026 | Jul 21, 2026 | Aug 13, 2026 |
| List price per 1M in / out | $1.50 / $9.00 | $1.50 / $7.50 | $0.75 / $3.75 to Dec 31 |
| Google's DeepSWE claim | Not published | 49.0% | 65.3% (v1.1) |
| Independent SWE-bench Verified | 78.8% | 79.6% ±1.80 | Not yet run |
| Our leaderboard status | Ranked | Ranked | Verifying |
Why does Flash keep shipping while Pro does not?
Google split the Gemini 3.5 generation into two launches back in the spring. Only one of them has ever arrived. Gemini 3.5 Flash went live at I/O on May 19 and did something Google had not managed before, with the cheap tier beating the previous flagship on coding. The Pro half has been in the wind ever since.
In July, Bloomberg reported that 3.5 Pro was months behind schedule, with coding the capability the team kept going back to fix, and quoted ten current and former employees describing a lab worried it had lost the frontier. Google shipped three models on July 21 and none of them was Pro; the official line was that Pro was "currently testing with partners," with no date. Bloomberg's story today, timed to this launch, says the delay persists and Google still would not give one.
- May 19, 2026Gemini 3.5 Flash launches at I/O 2026 The cheap tier beats Google's previous flagship on coding
- Jul 16, 2026Bloomberg reports 3.5 Pro is months late Coding is the capability the team keeps going back to fix
- Jul 21, 2026Three models ship, none of them Pro 3.6 Flash, 3.5 Flash-Lite and the locked-down Flash Cyber
- Jul 23, 2026vals.ai scores 3.6 Flash at 79.6% A statistical tie with 3.5 Flash despite the 12 point DeepSWE claim
- Aug 13, 2026Gemini 3.7 Flash ships at half price Bloomberg reports the Pro delay persists, still no date
- Jan 1, 2027Intro pricing ends 3.7 Flash doubles to $1.50 and $7.50 per 1M tokens
Who should actually switch today?
If you are already running 3.6 Flash in production, the arithmetic is easy. Same list availability, same API surface, half the bill for 140 days. Swap the model string, remove any minimal thinking level, and re-run your evals. The only thing to plan for is January 1, when the rate doubles back to 3.6 Flash's price. If your unit economics only work at $0.75, put a reminder in the calendar now rather than discovering it in a New Year invoice.
RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price
If you are on a rival model and shopping on price, be careful about reading the intro rate as Google's real number. OpenAI cut GPT-5.6 Luna to $0.20 and $1.20 per million at the end of July, and that model sits at 93.0% independently measured. Cheap and good are not the same axis, and neither is cheap and cheap-in-February.
If you are picking a model for long agentic runs, wait. The benchmarks Google leaned on today, AutomationBench and GDP.pdf in particular, are the ones with the least public track record, and a jump from 17.0% to 30.4% on a benchmark almost nobody has replicated is a weak basis for a production decision.
What it means for the market
Alphabet is doing something specific here, and it is visible in the pricing rather than the benchmarks. Halving the price of the volume tier while the flagship is stuck is a share play: Google is buying developer defaults on the model that actually carries the token volume, on the bet that whoever writes the integration in August is still there when Pro eventually lands. The revenue cost is real but bounded, since the rate reverts on January 1 and the cheapest customers are the ones least likely to churn over a doubling.
The signal for investors is not the DeepSWE number, it is the absence of a Pro date on the same day Google chose to make news. Two things are worth watching: whether Gemini 3.5 Pro gets a date before the end of the quarter, and whether the January 1 reversion actually holds or gets extended, which would tell you Google is defending share rather than running a promotion. This is analysis, not investment advice.
- The independent score. vals.ai ran 3.6 Flash two days after launch. Expect a neutral SWE-bench Verified figure for 3.7 Flash within the week, and expect it to land closer to 3.6 Flash's 79.6% than the DeepSWE jump implies.
- A Pro date, or another Flash. Google has now answered "where is Pro" three times by shipping something else. A fourth Flash release before any Pro date would say the flagship problem is structural, not schedule slip.
- Whether the intro price sticks. An extension past January 1 is a competitive response, not a promotion.
- Computer use leaving preview. It has been preview-tagged across three Flash releases now, and it is the feature that would actually differentiate this tier for agent builders.
Our take
Gemini 3.7 Flash is a good deal and a strange signal at once. The price is the most concrete thing Google shipped today, and for anyone already paying Google for tokens it is a straightforward win with a known expiry date. The benchmark story deserves the two day wait, because we watched the exact same shape of claim collapse into a statistical tie in July, and nothing about DeepSWE v1.1 makes it more comparable across labs than v1 was.
The uncomfortable part is what the version numbers now say out loud. Google's cheap workhorse is on 3.7. Google's flagship is on 3.1, because 3.5 Pro does not exist. A company that could ship its frontier model would have shipped it by now, and shipping a fourth Flash instead is starting to look less like a cadence and more like an answer.
- OfficialGoogle: Gemini 3.7 Flash, our most intelligent workhorse model Launch post with all five benchmark figures and the introductory pricing
- ReferenceGemini API model card: gemini-3.7-flash Token limits, modalities, tool support and the thinking-level restriction
- Benchmarkvals.ai: SWE-bench Verified The neutral bash-only harness that measured 3.6 Flash at 79.60% ±1.80
- DataGenZTech AI Coding Leaderboard Where 3.7 Flash sits unranked until an independent evaluation exists
Original analysis by GenZTech. Reporting on the launch and the continuing Gemini 3.5 Pro delay via 9to5Google.
