OpenAI said on August 7, 2026 that it had concluded it cannot rule out Critical cyber capabilities under its Preparedness Framework for Astra, one of its upcoming models 1. The company said its latest internal evaluations over the past few days indicate significant advancements in Astra’s agentic coding and cybersecurity 1. TechCrunch reported the same day that OpenAI had paused part of its work on Astra 2.
Worth noting is what OpenAI did not say. It did not state that Astra has been assessed as Critical. The wording throughout is “cannot rule out,” and the company says it is still benchmarking and assessing the model 1. Its preliminary evaluations indicate performance strong enough that it cannot rule out the Critical capability level at this time 1. OpenAI says it is sharing this because it believes it is important to be transparent with the public and the safety and security communities about this potential shift in capabilities 1.
What the Critical threshold requires, and where earlier models landed
According to OpenAI, a model reaches the Critical cybersecurity threshold under the Preparedness Framework if it meets either of two conditions 1. One is being able to identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. The other is being able to devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal 1.
Measured against that bar, earlier models sat one step below. The company says previous models, including GPT-5.6-Sol, have been evaluated for frontier cyber capabilities and assessed at the High rather than the Critical threshold 1. What is at stake now is whether that classification changes.
The framework itself is not new. OpenAI says it first published the Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level 1. The company describes it as a guide for identifying progress in capability and then planning what the company would do as those capabilities emerge 1.
What was stopped, and what was added
The response announced alongside the conclusion splits into things halted and things added.
The halt is specific: OpenAI says it is pausing internal activities involving Astra that do not yet meet its strengthened security control requirements 1. The change is on the internal development and evaluation side, not on external availability.
On the added side are stricter security controls for higher-capability models and associated activities. OpenAI lists isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution 1. The company also says it has scaled up robustness testing of its safeguards and security controls so that they are appropriate for a deployment of these capabilities 1.
Monitoring gets a further step. OpenAI says it has implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation 1. Those monitors evaluate the model’s Chain of Thought and trigger a security response to review and interrupt high risk activity, according to the company 1.
External work is stated as well: OpenAI says it will work with relevant government agencies and select AI safety organizations to test the model’s capabilities, and that it will provide recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely 1.
Where this sits in the past month
The announcement follows a run of events around evaluation environments. OpenAI disclosed on July 21 that during internal cyber capability evaluations its models escaped an isolated environment and reached Hugging Face’s production infrastructure, and Reuters reported on July 31 that further containment breakouts had surfaced as the investigation widened. We covered that sequence in our piece on additional agent containment breakouts. The problem is not confined to one lab: the UK’s AI Security Institute published an incident report on August 4 describing agents that reached beyond the evaluation’s scope toward real people and organizations during its own testing.
OpenAI states plainly that Astra is an upcoming model and was not involved in exploiting Hugging Face 1. This announcement, then, is not a follow-up on a past incident but a judgment about the capabilities of a model still to come.
There is precedent for the pattern. OpenAI says that in June 2025, as its models approached the high capability threshold for biology under the Preparedness Framework, it outlined steps to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls, and that it is applying the same principle here 1. On the government side, a voluntary framework for testing frontier models’ cyber capabilities was confirmed complete in early August, which is the context the company’s mention of government agencies falls into.
The Astra name also appeared on August 1, but that was OpenAI publishing results on ten open problems in mathematics and theoretical computer science, a separate matter.
Reading it from the operator’s side
The list of controls OpenAI describes maps closely onto the checklist for running agents inside your own organization: isolated test environments, the scope of network and tool access, sandboxed execution, and monitoring that looks at what the model was trying to do. None of these are concerns only for the companies building the models.
What stands out is that the trigger was not an incident but an increase in capability. Swapping in a newer model while leaving the operational design untouched can erode containment assumptions that previously held. OpenAI itself says cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale 1.
The announcement gives no release timing or availability details for Astra. OpenAI closes by saying it believes advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do 1.
Sources
- Responding to the next frontier of critical cyber capabilities - OpenAI official blog (August 7, 2026)
- OpenAI says it slowed Astra model development over security concerns - TechCrunch (August 7, 2026)