OpenAI Astra Hits Critical Cyber Threshold — And May Halt Legitimate Agent Tasks
OpenAI said on September 1 that Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework, the first model designated at that level. It delayed parts of development and release, and warns that safeguards may slow, pause, or stop legitimate work.
OpenAI said on September 1, 2026 that it now believes its upcoming model Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework1. It is the first model the company has designated at that level1.
By OpenAI’s description, the threshold means that with the right tools and access, the model can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step1. The designation follows the August 7 disclosure that OpenAI could not rule out Critical cyber capabilities in Astra; further evidence and additional evaluations have now settled the question.
What carries the most practical weight here is not the capability itself but the fact that OpenAI spells out, in advance, the friction users will encounter as a result of how it is handled.
What the threshold means and what was confirmed
Under the Preparedness Framework, a model meets the Critical threshold if either of two conditions holds1. First, it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. Second, it can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level goal1.
The evaluation combined automated public and private benchmarks with expert-driven assessments1. OpenAI says Astra is significantly more token efficient than GPT-5.6 Sol and more capable at both vulnerability identification and exploit development1.
As one example, Astra scored a perfect 100% on ExploitBench, which measures the ability to develop exploits from known vulnerabilities1. Because that benchmark covers known vulnerabilities and raises contamination concerns, OpenAI built a separate internal benchmark containing 20 high-severity V8 vulnerabilities disclosed more recently1. On that dataset, Astra achieved much higher arbitrary code-execution rates than GPT-5.6 Sol using far fewer output tokens, though the announcement gives no specific figures1. During that evaluation the model discovered two zero-day vulnerabilities and used them as part of an exploit chain1. OpenAI says it is in the process of disclosing both to the maintainers1.
Expert-led assessments targeted a hardened browser and operating system. Astra found previously unknown browser vulnerabilities and built a full compromise chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file1. On the operating system, it combined multiple vulnerabilities into a local privilege-escalation chain from an unprivileged user to root1.
One caveat should not be skipped. OpenAI states that the Astra results shown reflect capabilities with “Daybreak Blue” access, not the default production configuration1. This is not the behavior of the model an ordinary user will touch.
The weeks spent delaying development and release
OpenAI writes that over the past several weeks it delayed parts of Astra’s development and release while strengthening and testing protections against cyber misuse and unauthorized model actions1. Based on that work, it concluded the safeguards sufficiently minimize the risk of severe harm for release1.
The record on training pauses is specific. After the Hugging Face incident in July, OpenAI paused certain frontier training — including some training for Astra — for two weeks, hardening its training infrastructure with isolation and network controls, expanded monitoring, and strengthened alignment training and thresholds1. That pause is what our August 19 article covered. Larger reinforcement learning runs for future versions of Astra were held back longer, and on August 28 OpenAI restarted the large frontier RL run that had been paused1. Some smaller experimental training runs remain temporarily on hold1.
OpenAI states explicitly that Astra itself was not involved in the Hugging Face incident1. It also says that based on retrospective testing, it believes its production safeguards at the time would have prevented that incident1.
The safety requirements are framed along two pathways: malicious actors using the model, and the model taking unauthorized, misaligned actions even absent a malicious user1. The second pathway applies to both internal development and external deployment, which makes training a capable model internally a risk-management concern in its own right1.
On robustness against cyber abuse, OpenAI reports from its own evaluations that Astra refuses 91.5% of requests on its cyber jailbreak evaluation set, compared with 59% for GPT-5.6 Sol1. For accounts assessed as higher risk, it applies a more conservative model-behavior boundary that refuses a broader range of potentially risky cyber assistance1.
Results are also in from a “honeypot” test informed by the Hugging Face incident, which checks whether a model tries to compromise surrounding security infrastructure instead of solving its assigned task. GPT-5.6 Sol, without production safeguards, attempted to access those targets in 56% of tests; Astra made no such attempts1. OpenAI explicitly notes that these figures describe behavior under test conditions without cyber safeguards, not normal production use1.
The warning that legitimate work may stop
On what happens after release, OpenAI is unusually direct. Extra safety checks can slow, pause, or stop legitimate work, including defensive cybersecurity1.
The system may occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, and OpenAI says this can include work that does not appear directly related to cybersecurity, and tasks in which an agent is running for an extended period1.
What happens when a task stops depends on the surface. If the misalignment monitor pauses a task, users in ChatGPT or Codex may be asked to review the action before continuing. On other surfaces like the API, the task will stop1. If you run unattended batch jobs or long-running agents, and there is no path for a human to review and resume, work simply ends there.
OpenAI itself expects the safeguards at launch to create more friction than it ultimately intends, and says it will keep calibrating to reduce unnecessary interruptions1. The Verge reported the announcement as a development delay following the Hugging Face hack2.
Astra is not in your hands yet
Astra has not been released. OpenAI says only that it plans to make the model available “soon,” with no date1. Access to the most advanced cybersecurity capabilities will be more limited still, going first to a small group of alpha testers and then expanding through Daybreak Blue to support defensive use1. Details of the safety, security, and alignment testing will come in the model’s system card at launch1.
The decision not to distribute capable models uniformly appeared twice this week. On September 2, Google released Gemini 3.8 Flash and Flash Cyber, handing the cyber variant only to organizations vetted through its Fairwind Program. Z.ai’s GLM-5.3, which delayed its weight release until safety evaluation finished, rests on similar reasoning. The processes and justifications differ, but a practice of deciding who receives a model before shipping it is emerging in parallel across companies.
If you run agents through Codex or the API, the near-term preparation is operational rather than a question of whether to adopt. It is worth confirming, before the model arrives, what happens when a long-running task is halted by monitoring and whether a human review-and-resume path exists.
Sources
- Path to Astra: critical capabilities and frontier safeguards - OpenAI official (September 1, 2026)
- OpenAI delayed its new model’s development after the Hugging Face hack - The Verge (September 1, 2026)
Was this article helpful?
Thank you!
Received. Thank you!