3 min Applications

Gemini 4 “Argon” unveiled, but not yet released: what’s going on?

Gemini 4 “Argon” unveiled, but not yet released: what’s going on?

After nearly a year in which Anthropic and OpenAI steadily widened the gap with Google, Gemini is back in the game. Gemini 4 “Argon” is a model that, on paper, rivals Fable 5.1 and GPT-6 Astra. In some cases, it even outperforms them based on benchmark results. However, benchmarks don’t tell the whole story, the true test is still pending, as access remains highly restricted.

Google is restricting access to only participants in the security-driven “Fairwind Program,” its own equivalent of Anthropic’s Project Glasswing and OpenAI’s Daybreak. Nevertheless, its capabilities extend far beyond cybersecurity. For example, Gemini 4 is reportedly already helping Google optimize quantum computations, make optimal use of memory, and carry out large-scale code migrations.

Champion of complexity

Google is allowing the U.S. government to examine and review the model prior to its widespread release. Additionally, the company behind Gemini is testing the safety measures and gathering early feedback. It may therefore take some time before end users within organizations and consumers get their turn.

The pricing, however, is already known. It’s the same as for GPT-6.1 Sol: $2 per million input tokens and $10 per million output tokens. Cached input, which is cheaper for AI providers to process because the information is already stored on the servers, receives a 95 percent discount. Notably, the output limit increases from 64,000 to 1 million tokens.

From rust migrations to memory savings

Internally, Argon is already running on thousands of “Googlers’ ” machines. Agents are actively migrating C/C++ code to Rust, from libraries like re2 to the over 800,000 lines of the Zircon kernel in Fuchsia. In the video decoder libgav1, they replaced 32,000 lines of SIMD code. The result was a decoder that runs 2.7 times faster than the previous Rust port. Elsewhere, Argon agents freed up over 300 TiB of memory in Google’s data centers. This shows that Gemini 4 helps reduce costs in multiple ways, despite its undoubtedly significant inference costs.

In benchmarks, Google claims the top spot on DeepSWE v1.1 (77.9 percent), the Vals Index, and Zapier’s AutomationBench (51.3 percent). However, benchmarks, of all shapes and sizes, have become somewhat unreliable in recent months. This is because models such as Opus 5, for example, scored higher than Fable 5, even though every user had to conclude in the long run that the latter was significantly stronger.

Without cyber guardrails

For now, the focus remains on cybersecurity, despite Argon’s various strengths in other areas. The model can autonomously detect, validate, and patch vulnerabilities. Trusted defenders and internal Google teams receive the model without cyber guardrails. Wiz is already using Argon for its “Scan for Good” initiative, the company says. In doing so, the model found a critical vulnerability in healthcare software used by hospitals worldwide, a cyber threat that earlier frontier models had missed.

For the broader rollout, Google says it is strengthening safeguards against misuse for biological, chemical, radiological, and nuclear attacks, as well as against indirect prompt injection. In addition, it monitors the model’s chain of thought and actions for misalignment. Google is calling on the industry to maintain transparency in reasoning (a clear jab at Anthropic and OpenAI) so that models’ thought processes remain useful for detecting misalignment.

Read also: Google launches Gemini 3.1 Pro, an LLM for complex reasoning