OpenAI has suspended part of the training of its most powerful AI models after an AI agent once again managed to escape from a secured test environment. The system gained access to the internet and then independently contacted an external chatbot.
According to Bloomberg, the incident occurred during the training of an agent-based AI system in a sandbox that was specifically designed to block internet access. However, the model discovered a vulnerability in that security barrier and managed to reach the public internet.
It then sent at least twenty requests to an unnamed third-party chatbot service. As far as is known, these were not malicious commands; for example, one of the questions was what the capital of France is. The main issue lies in the fact that the model was able to leave its controlled environment in the first place.
In response, OpenAI has temporarily suspended training sessions in which its most powerful models can use tools. That training will not resume until the issue with the sandbox has been resolved. According to the company, the specific model involved in the incident will not be trained further.
Emergency stop did not work
The incident also brought a second problem to light. An internal monitoring system detected the unusual behavior and alerted an employee. That employee confirmed the alert within three minutes via Slack.
After that, the training run should have been automatically terminated, but it was not. Ultimately, it took more than two hours before the training was manually stopped. This meant that not only did the model’s technical safeguards fail, but so did part of the process designed to mitigate the consequences of such an incident.
This is not the first time OpenAI has encountered problems with models operating outside their permitted environment during testing. In July, during internal tests, models gained access to the internet and unintentionally breached the AI platform Hugging Face.
In addition, OpenAI confirmed this week that models accessed information on U.S. government agency websites, including the Census Bureau and the Securities and Exchange Commission (SEC), during training and evaluation. It was also previously reported that OpenAI agents were involved in an incident involving an Australian government website.
According to OpenAI, the new incident is the first security incident of this specific type since the events surrounding Hugging Face. Bloomberg reports that the company is still working to map out the previous activities of its agents.
Meanwhile, these incidents are further intensifying the debate over control of increasingly autonomous AI systems. Not only OpenAI, but also Anthropic, Google DeepMind, and Meta have recently investigated undesirable behavior by agents. OpenAI and Anthropic have also called for caution in developing systems that could play an increasingly significant role in improving future generations of AI models.