A Gemini 3.8 Flash model card briefly went live on Google DeepMind's own site within the past hour, then vanished, and it landed right as The Wall Street Journal reported Google could ship the coding-focused model as soon as today. Put those two things together and Google looks like it is one final review away from releasing a Gemini update built specifically to close the coding gap with Anthropic and OpenAI.

  • A page at deepmind.google/models/model-cards/gemini-3-8-flash/ went live, got spotted on Hacker News, and now returns a 404, the classic signature of a launch page published early by mistake.
  • The Wall Street Journal reports Google could release Gemini 3.8 Flash today, internally codenamed "Skimaki," with a focus on agentic coding.
  • Inside Google's Jetski developer platform, engineers reportedly preferred the new model's coding output to Anthropic's Claude Opus in head-to-head testing.
  • The predecessor, Gemini 3.7 Flash, currently sits 19th on GenZTech's independently-verified coding leaderboard at 80.8%, well behind Claude Opus 5's 97.0% and GPT-5.6 Sol's 96.2%, which is the gap 3.8 Flash is supposedly built to shrink.
Timeline of the Gemini 3.8 Flash leak Four-step sequence: a model card page appears on deepmind.google, it gets spotted on Hacker News, the page is pulled and returns 404, then the Wall Street Journal reports an imminent launch. STEP 1 Model card page appears at deepmind.google/.../gemini-3-8-flash/ STEP 2 Spotted on Hacker News, climbs to 22 points and 10 comments within the hour STEP 3 Page pulled, URL now returns a 404 Not Found STEP 4 WSJ reports a launch could land as soon as today, citing Jetski test results genztech.blog
Fig 1 The order events happened in, reconstructed from the Hacker News thread and WSJ's reporting.

What actually leaked, and when?

Someone at Google published the model card for Gemini 3.8 Flash to the public deepmind.google model-card directory before the model itself was ready to announce. That happens more often than launch teams would like: a CDN cache warms, a staging deploy targets the wrong environment, or a page goes live on a schedule that assumes an announcement will land first. Whatever the mechanism, a reader caught the URL, posted it to Hacker News, and it picked up traction fast, 22 points and 10 comments inside the hour, before someone at Google noticed and took the page down. Right now that URL returns a plain 404. Google has not commented on the page or confirmed a launch date.

RelatedGemini 3.7 Flash Ships at Half Price, 3.5 Pro Still Missing

Why does the leak line up with a Wall Street Journal report?

The timing is what makes this worth writing up rather than shrugging off as a routine leak. The same day the model card surfaced, the Journal reported that Google is preparing to ship Gemini 3.8 Flash, internally known as "Skimaki," with a specific target: agentic coding, the workflows where a model plans a change, edits multiple files, runs tests, and iterates without a human in the loop at every step. According to that reporting, Google ran the model through its internal Jetski developer platform and had engineers compare its output directly against Anthropic's Claude Opus, and the engineers reportedly preferred Gemini's answers. A smaller architecture than Google's flagship models also means it costs less compute to retrain, which the Journal's sourcing says let several research teams run reinforcement-learning passes on it in parallel rather than queuing for flagship-scale training runs.

Where does Google actually stand on coding right now?

Not close to the top, which is exactly why this release matters. On GenZTech's coding leaderboard, built from independent SWE-bench Verified runs rather than vendor claims, Gemini 3.7 Flash, the current Flash-tier model, sits 19th at 80.8%. Claude Opus 5 leads at 97.0% and GPT-5.6 Sol trails it by less than a point at 96.2%. That's not a close race; it's a 16-point gap between Google's best publicly-scored Flash model and the field's leaders.

ModelMakerSWE-bench VerifiedPricing (in/out per 1M tok)
Claude Opus 5Anthropic97.0%$5 / $25
GPT-5.6 SolOpenAI96.2%$5 / $30
Claude Sonnet 5Anthropic85.2%$2 / $10
Gemini 3.7 FlashGoogle DeepMind80.8%$0.75 / $3.75
Gemini 3.8 FlashGoogle DeepMindNot yet independently scoredUnconfirmed

Every number in that table except the last row comes from a neutral harness run, not a vendor's own press release, which is the whole reason the gap is worth taking seriously: Google's own launch posts for 3.6 Flash and 3.7 Flash claimed double-digit coding jumps that, once vals.ai actually ran the numbers, mostly turned out to be statistical ties with the model before it. If 3.8 Flash follows that same pattern, "preferred over Opus" in Google's internal testing could still land well short of Opus once someone outside Google measures it.

How did we get from 3.5 Flash to 3.8 in a matter of weeks?

  1. Jul 2026Gemini 3.5 Flash independently verified at 78.8% SWE-bench Verified
  2. Late Jul 2026Gemini 3.6 Flash claimed a 12-point DeepSWE jump; independent score landed at 79.6%, effectively a tie
  3. Aug 13, 2026Gemini 3.7 Flash 80.8% ±1.76 independently, the rare Flash release where the vendor's optimism actually held up
  4. Sep 2, 2026Gemini 3.8 Flash model card leaks page pulled within the hour, no official announcement yet
  5. Expected any hourPublic Gemini 3.8 Flash launch per WSJ sourcing
SWE-bench Verified score by Gemini Flash generation Bar chart showing independently-verified SWE-bench Verified scores rising from 78.8% for Gemini 3.5 Flash to 80.8% for Gemini 3.7 Flash, with Gemini 3.8 Flash shown as an unscored, dashed placeholder bar. 78.8% 3.5 Flash 79.6% 3.6 Flash 80.8% 3.7 Flash ? 3.8 Flash genztech.blog
Fig 2 · benchmark Independently-verified SWE-bench Verified scores across the Flash line. The dashed bar marks Gemini 3.8 Flash, which has no independent score yet.

What does this mean for Alphabet's AI story?

Coding is the use case investors currently watch most closely when they price AI spend, since it is the clearest example of a model doing billable engineering work rather than just answering questions. Google has spent the past several quarters telling that story with Gemini 3 Pro and the Deep Think tier, while the actual coding leaderboard has Google's models clustered well below Anthropic and OpenAI. A Flash-tier model that meaningfully narrows that gap, at a fraction of Opus or GPT-5.6 Sol's price, would give Google a genuine wedge into the developer-tooling market Anthropic and OpenAI have led since Claude Code and Codex took off. The signal for a reader tracking Alphabet is simple: watch the independent SWE-bench number the day it ships, not Google's own launch-post framing, and watch whether it prices meaningfully below Claude Opus while landing anywhere near it on real tasks. If it does, that is the combination that pulls price-sensitive engineering teams over. If it repeats the 3.6 Flash pattern of a big claimed jump that evaporates under independent testing, the market gap stays exactly where it was this morning.

RelatedDeepMind Disbanded the AlphaFold Team, Not AlphaFold

What happens next?

Two things to watch, and neither is complicated. First, whether Google actually ships today or the WSJ's timeline slips, since a leaked model card is evidence of intent, not a release. Second, and more important than anything in Google's own announcement, is what vals.ai's independent SWE-bench Verified run says once it evaluates the model, typically within a week or two of launch based on how the last three Flash releases played out. That number, not Google's internal Jetski comparison against Claude Opus, is what will actually tell you whether this closes the gap or just narrates one.

What to watch · Sep 2026
  • The official announcement. If Google ships today as WSJ's sourcing suggests, expect a same-day post on the Gemini blog and pricing details that will decide whether this actually undercuts Claude Opus on cost.
  • The independent score. vals.ai took 4 days to score 3.7 Flash and 3 weeks for 3.6 Flash. Whichever it is this time, that's the number that matters, not the internal Jetski comparison.
  • Whether "preferred over Opus" survives contact with a public benchmark. Google's last two Flash launches both claimed bigger coding gains than the independent runs confirmed.

Our take

A model card leaking an hour before a Wall Street Journal report names the same model is about as close to confirmed as an unannounced product gets, so treat the launch itself as close to certain. What's not certain is the headline claim. Google has now told this exact story twice, a Flash update that supposedly leapfrogs its predecessor on coding, and both times the number that actually got independently checked came in as a rounding error rather than a leap. "Engineers preferred it to Claude Opus" inside Google's own testing environment is a marketing data point, not a benchmark. The interesting question was never whether Gemini 3.8 Flash would ship today. It's whether Google finally publishes a coding claim that survives someone else running the eval.

Primary sources

Original analysis by GenZTech. Source reporting: Investing.com / Wall Street Journal.