New Google AI Model Said to Narrow Gap on Coding Ability
- Gammatek ISPL
- 6 days ago
- 5 min read
By Gammatek ISPL, Technology & Compliance Analyst at Gammatek ISPL
Published September 2, 2026 | 10 min read
Author block: Gammatek ISPL covers enterprise software decisions for Gammatek ISPL, where our own development and compliance-tooling teams evaluate AI coding assistants as part of building industrial safety and compliance software. This piece reflects our team's direct experience testing multiple AI coding tools, alongside reporting from the Wall Street Journal and Gizmodo (linked below). This is independent editorial analysis, not sponsored by Google or any AI vendor named.
Why This Matters Right Now
Google is reportedly on the verge of releasing a new AI model built specifically to close the gap on coding performance — and if you run or manage a software team, this isn't just tech-industry gossip. It signals a shift in how fast the tools your developers use every day are changing, and it raises a question most companies haven't actually answered yet: how do you evaluate an AI coding tool before betting your workflow on it? That's the gap this article actually closes — not just reporting the news, but giving you a framework to act on it.

What's Actually Being Reported
According to <cite index="9-1">an anonymously-sourced Wall Street Journal report</cite>, Google is preparing to release a new model — <cite index="9-1">referred to as Gemini 3.8 Flash, or internally as "Skimaki"</cite> — that insiders describe as having <cite index="9-1">substantially improved coding capabilities</cite>. Sources cited by the Journal say the release <cite index="9-1">could arrive within days of the report surfacing</cite>.
It's worth being precise about what this model actually is. <cite index="9-2">This won't be a massive, frontier-scale model in the way OpenAI's still-unreleased "Astra" is being discussed</cite> — <cite index="9-2">Flash-tier Gemini models are built to run faster and more cheaply than Google's largest models</cite>, which matters for a very practical reason: <cite index="9-2">tools that aggregate and route between different AI models based on cost and speed have become a growing part of how companies actually deploy
AI in production</cite>. A cheaper, faster coding model that performs well isn't a side release — it's aimed squarely at how AI gets used day-to-day inside real engineering teams, not just in benchmark demos.
The timing also isn't incidental. <cite index="9-3">Google DeepMind's longtime CEO, Demis Hassabis, recently stepped aside from the CEO role</cite>, and <cite index="9-3">Google co-founder Sergey Brin has reportedly been taking a more hands-on role in AI strategy, pushing for faster shipping of competitive Gemini models</cite>. <cite index="9-4">Hassabis's new focus is described as centered on scientific research rather than day-to-day model releases</cite> — a notable shift for a leader who <cite index="9-4">had previously signaled caution about rushing frontier model releases given the risks involved</cite>.
Perhaps the most concrete detail in the report: <cite index="9-5">Google's own insiders reportedly tested the new model on an internal coding tool called "Jetski," and engineers involved said they preferred it over Anthropic's Claude Opus</cite>. That's an informal, anonymous claim — not a published benchmark — but it's the kind of internal signal that tends to precede a real product push.
Original Analysis: What This Actually Means for Software Teams
Here's the part most coverage of stories like this skips: the news itself — "Google has a new coding model" — is almost never the actionable part. Model releases happen constantly now. What actually matters for a software or compliance-tooling team is how you decide whether to adopt it, and most companies don't have a real process for that. They either switch tools reactively whenever a new headline drops, or they stick rigidly with one vendor regardless of whether it's still the best fit.
At Gammatek, our development team maintains and updates industrial compliance and safety software — which means our code has to be reliable, auditable, and consistent, not just fast to produce. When we evaluate any AI coding tool (including previous Gemini releases, Claude models, and GitHub Copilot), we don't just look at raw coding benchmark scores. We look at:
Consistency across runs — does the model produce similar-quality output on the same type of task repeatedly, or does quality vary widely?
Behavior on legacy or unusual codebases — compliance software often touches older systems and integrations; a model that excels on greenfield code isn't automatically useful here.
Explainability of suggestions — for regulated-industry software, being able to explain why a piece of code does what it does matters for audit purposes, not just whether it runs correctly.
Cost at scale — a Flash-tier, faster/cheaper model matters more in practice than a slower flagship model for day-to-day use, which is exactly the lane this rumored Gemini release is aimed at.
A simple comparison framework, based on how we evaluate new coding models internally:
Evaluation criteria | Why it matters | What to actually test |
Benchmark performance | Useful as a filter, not a decision | Check if benchmarks reflect your actual codebase type (web, embedded, data pipelines) |
Cost per task at scale | Flash/lightweight models often win here | Run a real sprint's worth of tasks through both models, compare token cost |
Consistency | Benchmarks test peak performance, not average reliability | Run the same prompt 5-10 times, check variance in output quality |
Legacy code handling | Most real companies aren't working in greenfield code | Test against an actual older module, not a fresh demo repo |
Audit/explainability | Critical for regulated industries | Ask the model to explain its own suggestion — can a reviewer understand the reasoning? |
This is the piece missing from most "new AI model" coverage: a new model narrowing a benchmark gap is interesting, but it doesn't tell you whether it's the right tool for your team's actual codebase, compliance requirements, or workflow. The gap between "wins a benchmark" and "should be part of your production workflow" is exactly where most companies get this decision wrong.
The Bigger Pattern This Fits Into
This rumored Gemini release fits a broader pattern that's been building through 2026: coding has become the primary battleground for major AI labs, more than general chat or reasoning benchmarks. OpenAI reportedly shifted focus toward business and productivity tooling earlier in the year, and multiple labs have released models specifically optimized for software engineering tasks rather than general-purpose use. The signal here isn't "Google has a new model" — it's that the entire competitive landscape has organized itself around coding as the proving ground, which means the tools available to your engineering team will keep changing faster than most companies' internal evaluation processes can keep up with.
For any company running regulated or safety-critical software — industrial compliance platforms very much included — that pace creates a real operational question: how do you stay current on tooling without introducing risk by adopting unproven models into production systems? That's less a story about Google specifically and more a permanent feature of how software will get built from here on.
What To Actually Do With This
If you manage a development or technical compliance team, the practical takeaway isn't "switch to Skimaki when it launches." It's:
Build a lightweight, repeatable evaluation process (like the framework above) so you're not making ad-hoc decisions every time a new model drops.
Separate "interesting benchmark" from "production-ready for us" — these are different questions with different evidence requirements.
For regulated or audit-sensitive codebases specifically, weight explainability and consistency more heavily than raw speed or cost, even when a new model looks attractive on paper.
How This Connects to Compliance-Critical Software
AI coding tools are moving fast — but for companies building or maintaining software that supports regulatory compliance (safety audits, EHS reporting, permit-to-work systems), the tooling decision is only half the picture. The bigger question is whether your underlying compliance processes are documented and auditable regardless of which coding tools your team uses this quarter.
[See how Gammatek's compliance platform supports audit-ready documentation across your software and operational processes →




Comments