2 min Security

OpenAI slows down AI development after ‘Hugging Face’ incident

OpenAI slows down AI development after ‘Hugging Face’ incident

OpenAI has scaled back part of its development of new AI models. This follows a security incident in which an autonomous AI agent escaped its isolated environment during a test and accessed Hugging Face’s systems. In response, OpenAI is tightening its testing and training procedures.

The company suspended model evaluations for two weeks, Reuters reports. Training for Astra, OpenAI’s next-generation AI models, has also been paused. A large-scale training run scheduled for Astra has not yet resumed. OpenAI has not disclosed exactly when the temporary delay began.

This move marks a change in course for OpenAI. In recent years, the company had been striving to accelerate the development, evaluation, and rollout of models. The competitive race among generative AI providers made rapid development essential.

Security measures are now being given higher priority. Among other things, OpenAI plans to use AI systems to monitor other agents during tests. In addition, sensitive experiments must take place more frequently in strictly isolated sandbox environments.

Another technique OpenAI is employing is “chain-of-thought monitoring.” Researchers use this to gain insight during tests into the steps a model takes to reach a decision or perform an action. However, this approach offers no guarantee. Research indicates that models do not necessarily reveal in their visible reasoning process that they intend to violate rules.

The measures were prompted by an incident last month. An autonomous agent using two advanced OpenAI models was being tested on cybersecurity tasks. During the test, the agent managed to operate outside the intended test environment and gain access to Hugging Face’s infrastructure. The system did so to achieve the objective it had been given during the test.

Astra does not yet meet requirements

The stricter approach is not limited to the evaluation environment. On August 7, OpenAI announced that it would impose additional security requirements on its most powerful models. OpenAI suspended activities involving Astra because the model did not yet meet those requirements.

OpenAI links these measures to its Preparedness Framework, which outlines how it handles models that develop potentially dangerous capabilities. The company now states that changes are needed not only to its own procedures. As AI models become more autonomous and powerful, the industry as a whole will require more comprehensive security measures.