During its investigation into the widely discussed hack at Hugging Face, OpenAI has found evidence that other autonomous AI agents have also ventured outside their secured test environment. These are reportedly multiple incidents that are now being re-examined. The discoveries are increasing pressure on AI companies to implement better security measures.
According to Reuters, the new cases came to light during the analysis of log files following the Hugging Face breach. In that incident, an AI agent managed to escape from a controlled test environment earlier this month. The agent then remained active within Hugging Face’s network for several days. OpenAI had previously stated that accounts belonging to four other organizations were also compromised in the incident.
The investigation now reveals that more similar incidents have surfaced. An OpenAI spokesperson referred to an earlier statement in which the company indicated that it was investigating more than just the Hugging Face attack. The company is also investigating “broader activities of our models.” According to sources, the additional escapes were limited. Furthermore, there are no indications that the AI agents involved left OpenAI’s internal network.
Anthropic is also investigating previous incidents
The expansion of the investigation comes shortly after Anthropic announced that AI models were responsible for multiple digital breaches at other organizations. These reportedly involve three incidents dating back to April. This suggests that both leading AI laboratories are struggling to control AI systems that are operating with increasing autonomy.
This is fueling concerns among security researchers. Maurice Chiodo, a mathematician at the University of Cambridge’s Center for the Study of Existential Risk, sees a pattern: the development of powerful AI agents is outpacing security measures. “An industry is emerging in which the people building these systems are struggling to keep them under safe and responsible control,” he says.
A key point of criticism is that the incidents only came to light after the fact. Techzine previously reported that OpenAI only became aware of the Hugging Face breach after the affected company had stopped the attack. The company then notified the FBI and made the incident public. OpenAI has partially disputed that account. The company did not specify which parts are inaccurate.
Anthropic also acknowledged that real-time monitoring likely could have detected the problems sooner. According to the company, monitoring was available. However, it was not applied to this threat scenario due to a misunderstanding with an external partner.
Increasing political pressure
The new revelations are drawing political attention. The European Commission confirmed on Friday that it had held discussions with both OpenAI and Anthropic regarding the recent incidents. In the United States, President Donald Trump said the government is considering additional measures regarding the development of advanced AI systems.
Senator Mark Warner, vice chairman of the Senate Intelligence Committee, also views the events as a case for stricter regulation. According to him, the incidents underscore the need for advanced AI models to undergo comprehensive safety testing. This must occur before they are deployed on a large scale.