3 min Security

Gemini hacks three companies during security test

Gemini hacks three companies during security test

A Google AI model infiltrated three companies during a security test. Gemini unintentionally gained access to the internet and managed to reach external systems. According to Google, the model stopped itself each time it discovered it had accessed environments belonging to other organizations.

The incidents occurred in May during a test conducted by security firm Irregular, according to The Wall Street Journal. Gemini participated in a so-called “capture-the-flag” exercise, in which it was tasked with retrieving information from a fictional company. However, that company shared the same name as an existing business. Due to an error in the test setup, the AI agents were also able to access the public internet, even though that was not the intention.

In one instance, Gemini tried various passwords until it gained access to a secure system. During two other test runs, the model searched the internet for information about the company. In doing so, it found login credentials for other companies in public repositories. Gemini then used that information to gain access to their systems.

According to Google, in all cases the model stopped as soon as it realized it had accessed other organizations’ systems. The three affected companies have been notified, as have U.S. federal authorities. Google will not disclose which version of Gemini was used in the tests. The company does state, however, that it was not its latest model.

Google does not view this as a model going off the rails

Google does not consider the incidents to be cases of so-called “model misalignment,” in which an AI model acts contrary to the intentions or values of its developers. Instead, the company points out that the safety measures worked: Gemini ultimately recognized the error and ceased its actions.

That explanation is not without controversy. Jack Cable, CEO of AI security firm Corridor, told The Wall Street Journal that what is particularly relevant is that an AI agent operated independently beyond the scope of its mandate and managed to infiltrate other organizations’ systems.

Google was informed of the incidents by Irregular in late July but did not disclose them publicly on its own. The company confirmed the hacks after The Wall Street Journal asked about them.

Previous incidents involving AI agents

This incident is not an isolated one. Irregular was also involved in similar tests with models from OpenAI, Anthropic, and Meta, in which AI agents similarly managed to access systems belonging to other organizations.

Irregular told CNBC that the same technical flaw also played a role in those earlier incidents. In response to the events, Google has modified its testing process in collaboration with Irregular.

The incidents raise the question of how AI agents can be tested safely when they are capable of using tools independently, searching the internet, and detecting security vulnerabilities. Another question is when AI companies should disclose such incidents.