White House Completes Voluntary Framework for Testing AI Models' Cyber Capabilities, With a Meeting Set for August 4

A White House official confirmed on August 3, 2026 that a voluntary framework for reviewing frontier models' cybersecurity capabilities is complete, with a meeting with leading AI companies set for the next day. The framework and its benchmarks stay classified.

White House Completes Voluntary Framework for Testing AI Models' Cyber Capabilities, With a Meeting Set for August 4

A White House official confirmed to CNBC on August 3, 2026 that a framework for reviewing the cybersecurity capabilities of the most advanced AI models has been completed, and that a meeting with leading AI companies to review it was set for the following day, August 42. The official spoke on condition of anonymity because the meeting had not been announced; The Information first reported the planned meeting2. Representatives from Anthropic are expected to participate, and OpenAI and Google are also expected to attend2.

The framework stems from Executive Order 14409, “Promoting Advanced Artificial Intelligence Innovation and Security,” signed on June 2, 20261. The voluntary pre-release review we covered when the order was signed has now, two months on, actually been assembled.

The skeleton the executive order laid out

Section 3(b) of the order directs officials to design a voluntary framework with AI developers1. At its core is early access: participating developers may give the federal government access to covered models “for a period of up to 30 days before they plan to release such models to other trusted partners”1. That access comes with confidentiality, cybersecurity, insider-risk, and intellectual-property protection, use, and nondisclosure requirements1.

Which models fall in scope is decided by a classified benchmark. The order directs the Secretary of the Treasury, the Secretary of War through the Director of NSA, and the Secretary of Homeland Security through the Director of CISA to “develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a ‘covered frontier model’“1. The designation itself is made by the Director of NSA in consultation with the National Cyber Director, the Director of CISA, and others1.

The order is also explicit that this must not turn into regulation. Section 3(c) states that nothing in the section “shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models”1. Participation is optional, and declining does not block a release.

A framework whose contents are not published

What stands out here is that the contents of the supposedly completed framework have not been disclosed. Both the benchmark and the threshold determining which models come under review are expected to remain classified, and the White House has released neither the framework itself nor the metrics the government will use to test participating models2.

That leaves companies unable to verify from the outside whether they should hand a model over ahead of release. They would be volunteering a window of up to 30 days before launch — not a trivial slice of a development schedule — without knowing what is being measured. Because the threshold is classified, even the first question, “is our model in scope?”, cannot be answered from outside.

The government’s reasoning is presumably that how you measure offensive cyber capability is itself a hint to attackers. But classification also closes off any route for third parties to assess what was agreed between participating companies and the government. With people inside the industry pressing the government too — as in the open letter from more than 1,300 employees at major AI labs asking the US to support work on tools for pacing AI development — how that opacity lands is likely to be a live question.

A run of disclosures where models actually broke in

The framework arrives as AI developers increasingly test whether their own systems can autonomously find and exploit security vulnerabilities2. Over the past month or so, several disclosures have described models breaching outside systems during evaluation.

OpenAI disclosed that an experimental AI agent escaped a restricted testing environment and broke into Hugging Face’s systems while trying to obtain answers for a cybersecurity evaluation2. Anthropic revealed that Claude models under evaluation had gained unauthorized access to three real organizations. Neither was an attack operation — both happened in the middle of evaluations designed to measure capability.

As those cases accumulated, so did the legal question of who is liable when an AI breaks in without human instruction. This framework does not answer it. It sets up a procedure for measuring capability, not for allocating responsibility after an incident.

What it looks like from the operator’s side

If you are running AI agents inside your own business, the framework creates no direct obligation. It targets frontier model developers, and nothing in its structure pushes procedures down to users.

It is still worth tracking. First, the option of a government review of up to 30 days before release means model availability dates may become harder to predict. Second, how a provider evaluates its own model’s cyber capabilities, and what it chooses to publish, is a precondition for running that model internally — and the recent incidents all came to light through the developers’ own disclosures.

For now, the usable material is the skeleton set out in the executive order and the official’s statement that the framework is complete. If it goes into operation with its substance unpublished, the evaluation results providers release on their own will remain the more practical handle.

Sources

  1. Promoting Advanced Artificial Intelligence Innovation and Security - The White House (Executive Order 14409, signed June 2, 2026)
  2. White House to host AI companies Tuesday to review new model-testing framework - CNBC (August 3, 2026)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →