2 min Security

OpenAI tightens security for future Astra model

OpenAI tightens security for future Astra model

OpenAI has implemented additional security measures for Astra, an AI model yet to be released. According to the company, recent internal tests have demonstrated such strong cyber capabilities that it cannot be ruled out that Astra will reach the highest risk threshold in its own Preparedness Framework.

According to OpenAI, evaluations over the past few days indicate significant progress in agentic coding and cybersecurity. Earlier models, including GPT-5.6-Sol, were still classified as “High” in this area. Based on preliminary results, Astra may fall into the higher “Critical” category.

OpenAI emphasizes that it has not yet been determined whether Astra actually meets those criteria. Benchmarks and further evaluations are still ongoing. The “Critical” threshold includes, among other things, independently discovering and developing functional zero-day exploits for well-secured systems or autonomously devising and executing new attack methods against heavily secured targets.

OpenAI also explicitly states that Astra was not involved in the previously reported incident involving Hugging Face.

The announcement follows several recent incidents during cybersecurity tests involving advanced AI models. These include the Hugging Face incident and similar occurrences during tests of Anthropic’s models.

Activities suspended

In light of the test results, OpenAI is tightening security measures surrounding Astra’s further development. These measures include isolated test environments, restrictions on network and tool access, enhanced protection and encryption of model weights, additional monitoring, and execution within sandbox environments.

Internal activities involving Astra that do not yet meet these stricter security requirements have been suspended for the time being. OpenAI is also implementing monitoring of all agentic applications of Astra during training and evaluation. This involves checking for high-risk actions and potential misalignment. In the event of a high risk, a security procedure can have the activity assessed and interrupted.

OpenAI also plans to further test Astra’s capabilities in collaboration with relevant government agencies and selected AI security organizations. External testing partners will receive additional security guidelines for higher-risk evaluations and workloads.

The company says it ultimately intends to deploy advanced cyber models for defensive purposes as well, for example, to detect and remediate vulnerabilities before attackers can exploit them.