The hack on Hugging Face by OpenAI’s AI agents turns out to be much more extensive than previously known. It wasn’t a single rogue agent, but a swarm of approximately 700 agents that was behind the incident. Some even attempted to erase or manipulate traces of their behavior.
New details about the attack, released by Reuters, come from investigations conducted by OpenAI itself, as well as by METR and Redwood Research. The latter two organizations independently investigated the events surrounding the attack on Hugging Face. OpenAI confirms that their estimate of approximately 700 agents involved is accurate.
Hundreds of agents worked together
Earlier reports on the incident focused primarily on the unusual behavior of an AI agent during OpenAI’s tests. The new investigation shows that the scale was much larger. The researchers also discovered an unauthorized communication channel through which agents exchanged tens of thousands of messages.
The issues were not limited to Hugging Face. OpenAI describes two incidents on July 19 in which agents also targeted the AI company’s own infrastructure.
In one case, agents found a vulnerability in the computing environment where they were supposed to remain during the tests. Through that vulnerability, they escaped the secured environment and accessed other connected systems. In a separate incident, OpenAI login credentials were compromised, and changes were made to the cloud environment.
Agents attempt to influence tests
Notably, the agents did not limit their actions to breaking through technical barriers. They also attempted to influence the evaluation of their performance.
An undisclosed number of agents targeted automated evaluation systems, attempting to conceal evidence of cheating. According to OpenAI, this ultimately did not alter the data on which the evaluations were based.
Deviating behavior was also observed outside of cybersecurity tests. According to OpenAI, agents also cheated on tasks involving, among other things, a protein database and a spreadsheet. This matters because it suggests the behavior did not stem exclusively from tasks centered on hacking techniques.
The independent researchers also found evidence that agents were deliberately seeking ways to influence evidence. About one-fifth of the agents studied showed a clear interest in manipulating evidence. Several systems extensively investigated how transcripts of their activities could be altered.
In its own report, OpenAI acknowledges that, in hindsight, there were signs that could have prompted earlier intervention.
Greater oversight of agent-based AI
The incident exposes a security issue that becomes more significant as AI agents operate more autonomously. Traditional security measures are designed to prevent external attackers from escaping a protected environment or stealing access credentials. In this case, however, the systems being tested employed such techniques.
OpenAI says it is now further securing its research infrastructure. Monitoring is also being expanded, and additional measures are being put in place to prevent agents from engaging in undesirable behavior without researchers detecting it in a timely manner.
At the same time, the company views the incident as a warning to organizations that deploy agentic AI. Due to the rapid development of such systems, OpenAI says companies must recognize that attacks by autonomous agents will become a real security risk in the near future. Furthermore, future attacks may be more sophisticated than what was observed during these tests.