Google unveils Gemini 4 Argon, limited to cyber defenders for now
Google says Gemini 4 Argon tops GPT-6 Astra and Claude Opus 5.5 on 13 of 19 published benchmarks and can write 1 million tokens at once. Only vetted cyber defenders can use it today.
Google DeepMind announced Gemini 4 Argon on Wednesday, calling it a new frontier model for coding, enterprise knowledge work and cyber defense. It is rolling out first to vetted cyber defenders through what Google calls its Fairwind Program, with developers, enterprises and consumers to follow. Google gave no date for that wider release and said it is still strengthening safeguards.
https://x.com/GoogleDeepMind/status/2105388084154056939
Google says the model scores 77.9% on DeepSWE v1.1, a test of long software engineering tasks, and that its output limit has risen from 64,000 tokens to 1 million. TheRundown AI, summarising Google's own tables, counted first place on 13 of the 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5. These are Google's numbers, chosen by Google.
Outside measurements arrived within hours. Artificial Analysis said Argon scores 53 on its Intelligence Index, level with GPT-6 Astra and one point ahead of GPT-6.1 Sol, and that hallucinations are lower. Arena said the model reached first place in its Text Arena with 1,525 points, and eighth in its WebDev coding board. Argon trails Claude Sonnet 5.5 and Opus 5.5 on Terminal Bench 4, Artificial Analysis said.
https://x.com/ArtificialAnlys/status/2105392625788637299
Price is the sharper claim. Google lists introductory rates of $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20, with a 95% discount on cached input. At the discounted rate Artificial Analysis put the cost of a task at $1.99, against $3.26 for Astra. At standard prices it put Argon at $3.98 a task.
Not everyone inside Google is convinced. Bloomberg reported, citing employees, that the model performs well on benchmarks but struggles with some real-world coding, including front-end design. Google told Bloomberg it would be inaccurate to say Gemini 4 underperforms in coding. Because access is restricted, few outsiders can test the model, and none of the benchmark claims have been independently reproduced yet.