Google has unveiled Gemini 4 Argon, and the first people getting their hands on it will be looking for security holes.
The new model, Google’s first flagship generation since Gemini 3 last November, is initially going to a small group of cybersecurity defenders. Paid API customers and Google AI Ultra subscribers are next, with broader access dependent on further testing.
Google says Argon is designed for longer, more demanding jobs across software engineering, cybersecurity, legal work and finance. For founders already using AI to build products or tackle operational work, that’s an appealing pitch. Whether it holds up outside a benchmark is the more useful question.
The early picture includes some impressive claims, alongside reported doubts within Google itself.
Security teams get the first look
Access starts through Google’s Fairwind Program for trusted cyber defenders. Those participants will be able to use Argon without its usual cyber guardrails, giving them access to its full security capabilities, according to Google.
Security firm Wiz is already using the model through its Scan for Good initiative. Google says Argon identified a critical vulnerability in healthcare software used by hospitals worldwide, which earlier frontier models had missed.
That would be a meaningful result, although it remains a company-reported finding rather than a broad measure of how reliably Argon spots vulnerabilities.
Google is also participating in the US government’s voluntary process for giving officials access to models before release.
Bigger outputs and ambitious performance claims
Google says Argon scored 77.9% on DeepSWE v1.1, a benchmark for extended software engineering tasks. It also tied for first place on CWE-bench, which tests a model’s ability to fix security flaws, with a score of 68%.
Axios reported that Argon outperformed OpenAI’s GPT-6 Astra on several coding and knowledge-work benchmarks.
One particularly substantial change is the output limit, which rises from 64,000 tokens to one million. That gives the model considerably more room to produce lengthy responses, although more output doesn’t automatically mean more useful work.
Google says thousands of its employees already use Argon. It also credits teams of Argon agents with freeing more than 300 TiB of memory across its data centres.
Introductory API pricing is set at $2 per million input tokens and $10 per million output tokens. For anyone building a product around it, the eventual cost will depend on how much the model needs to generate to finish a job successfully.
Inside Google, the verdict reportedly varies
Bloomberg reported that some employees have found a gap between Argon’s benchmark performance and its usefulness in everyday work, including difficulties with certain coding tasks.
The outlet also reported that Google dropped Gemini 3.5 Pro, which it had previously promised for June.
Google pushed back on the suggestion that Gemini 4 underperforms in coding. One employee told Bloomberg there was broad agreement internally that the model belongs among the leading systems.
Those accounts leave a familiar problem for prospective users: strong test results can help identify a model worth trying, but they don’t tell you how much checking and fixing your particular workload will require.
What founders should watch
If Argon can reliably handle longer engineering tasks, it could be useful for small teams with more work than people. Its security capabilities may also prove valuable, especially if the early vulnerability-finding results translate into repeatable performance.
For now, most founders will have to wait for access. When it arrives, the useful test will be a real task from their own backlog, with the time spent reviewing and correcting the result counted too.
A model that produces more code is easy enough to get excited about. One that consistently leaves you with less work would be worth paying for.