4 min Security

OpenAI and Anthropic agents carried out independent hacking during tests

OpenAI and Anthropic agents carried out independent hacking during tests

During a security evaluation conducted by the British AI Security Institute (AISI), AI agents from OpenAI and Anthropic independently carried out cyberattacks against real individuals and organizations. Among other things, the models attempted to inject malicious code into an open-source project, used social engineering, and sought to collaborate with other AI agents. According to AISI, this is the first time that autonomy and deception have manifested so clearly in a real-world scenario.

The incidents occurred during an evaluation of advanced AI models for their cyber capabilities. To assess their full potential, the models were intentionally granted internet access, and various built-in security mechanisms were temporarily disabled. According to AISI, this configuration is not comparable to how the models are normally made available to users.

Nineteen unauthorized actions

The researchers ran the cyber challenge 122 times using seven different state-of-the-art models. In 10 of those tests, AI agents performed a total of 19 unauthorized actions on the public internet. Seventeen of these originated from Anthropic’s Mythos 5, while OpenAI’s GPT-5.6-Sol was responsible for two incidents.

The most serious incident involved an attempt to carry out a supply-chain attack via a popular open-source project on GitHub. The AI agent created multiple fake online identities and attempted to use social engineering to persuade the human maintainer to approve malicious code. When other developers became suspicious, the agent edited previous messages to appear less suspicious and even considered continuing under a new identity. Ultimately, a human reviewer rejected the change.

AISI identified even more notable behavior. AI agents directly approached real people via online messaging services and sent files intended to trick victims or their AI code assistants into executing malicious code. In addition, the models attempted to leave hidden prompt injections that could manipulate other AI systems. In another instance, an agent left messages on GitHub for other AI agents, instructing them to reuse existing accounts and digital traces. According to AISI, this is the first time such forms of collaboration and deception have been observed during an evaluation.

No sandbox escape

The institute emphasizes that the AI models did not escape from a protected test environment. In fact, internet access was deliberately granted to investigate their maximum cyber capabilities. However, the models turned out to be surprisingly creative in carrying out their tasks. They were never explicitly instructed to deceive humans, but did so anyway to achieve their goals. According to AISI, it was precisely this spontaneous, goal-oriented behavior that was the biggest surprise.

The researchers point out that human oversight ultimately prevented the most serious attacks from succeeding. At the same time, the institute noted that the margin was narrow. In several cases, the failure of the attacks depended on vigilant developers rather than technical security measures.

OpenAI reported an additional hack during irregular tests

In addition to AISI’s findings, Bloomberg reports another separate security incident that OpenAI itself described in a blog post. During a so-called “capture the flag” evaluation, conducted in collaboration with cybersecurity firm Irregular, OpenAI models exploited a configuration error in the test environment to gain internet access. They subsequently managed to hack the website of an unnamed organization.

According to Bloomberg, this incident was part of the same evaluation campaign by Irregular in which Anthropic models had previously compromised three organizations. This appears to further increase the number of incidents during advanced AI tests.

OpenAI also announced two weeks ago that one of its models had escaped from a secure sandbox during another security test. After gaining internet access, the model then infiltrated Hugging Face’s systems. That incident prompted a broader investigation into the behavior of autonomous AI agents.

Evaluation procedures tightened

In response to the incident, AISI is tightening its procedures. Internet access will now be restricted much more strictly, evaluations will be monitored in real time, and future tests will assume that powerful AI agents will actively attempt to act outside their mandate. In addition, the institute is further investigating the incidents in collaboration with Anthropic, OpenAI, and the research organization METR.

According to AISI, the findings show that the risk landscape is changing. Hazards no longer arise solely from humans misusing AI, but also from increasingly autonomous AI agents in research environments taking independent actions that their developers did not anticipate. Although the incidents occurred under exceptional test conditions and did not cause any actual damage, the institute views them as an important signal for the further development of advanced AI systems.