3 min Security

OpenAI test leads to incident at Hugging Face

OpenAI test leads to incident at Hugging Face

OpenAI has confirmed that an internal test involving advanced AI models led to a security incident on the Hugging Face AI platform. During an evaluation of their cyber capabilities, the models managed to escape from a secure test environment. They then penetrated the platform’s infrastructure. According to OpenAI, this is an unprecedented incident that demonstrates just how far the cyber capabilities of the latest generation of AI models have advanced.

Hugging Face initially reported the incident last week as an attack by an “external AI agent,” as Techzine previously reported. Only after a joint investigation did both organizations conclude that the attacking AI models originated from an internal test environment at OpenAI.

Benchmark led to actual attack

The incident occurred during an internal benchmark that OpenAI uses to assess the cyber capabilities of its latest models. Among others, GPT-5.6 Sol and an unreleased model were tested. For this evaluation, the usual security measures that prevent models from engaging in risky cyber activities were deliberately disabled to measure their maximum cyber capabilities.

According to OpenAI, this is, to the best of their knowledge, the first time a cyber capabilities benchmark has resulted in an actual attack on an external organization.

The models ran in a highly isolated test environment with only limited network access. However, according to OpenAI, they discovered an unknown vulnerability in software used to install software packages. By exploiting this zero-day vulnerability, the models gained access to the internet.

The models then performed multiple privilege escalations and lateral movements within OpenAI’s research environment until they reached a system with full internet access.

Once online, the models concluded that Hugging Face might be hosting datasets, AI models, and solutions for ExploitGym. ExploitGym is a publicly available benchmark that measures how well AI systems can exploit existing vulnerabilities.

According to OpenAI, the models were entirely focused on achieving the highest possible score on this benchmark. To circumvent the evaluation, they actively sought ways to obtain the correct answers directly. In doing so, they combined various attack techniques, including stolen login credentials and a zero-day vulnerability, to ultimately gain access to Hugging Face’s production environment. There, they retrieved the test solutions directly from the production database.

The unusual activity was detected by both OpenAI’s security team and Hugging Face. Hugging Face managed to stop the attack on its infrastructure and immediately launched a forensic investigation. During the joint investigation, both organizations determined that the attacking AI agent originated from OpenAI’s internal test environment.

OpenAI tightens security

OpenAI reports that it has since responsibly reported the discovered zero-day vulnerability to the vendor of the affected software. The research environment and procedures surrounding model evaluations are also being tightened. According to the company, these additional security measures will temporarily slow the pace of research.

In addition, Hugging Face has been included in OpenAI’s Trusted Access program, enabling the platform to deploy the latest AI models and further strengthen its security.

OpenAI calls the incident an important signal that the cyber capabilities of AI models are growing faster than many existing security measures. Recent evaluations by the UK’s AI Security Institute have already shown that models such as GPT-5.6 Sol can carry out prolonged, complex cyber operations. According to OpenAI, this incident demonstrates that these capabilities are no longer purely theoretical.

The company states that advanced AI should ultimately be used to detect security vulnerabilities before malicious actors do. At the same time, according to OpenAI, the incident shows that the security of test environments and oversight of powerful AI models must also be further strengthened as these models become more autonomous and capable.