Cover photo

The US Finalized AI Cyber Tests Without Publishing the Test

A voluntary framework now exists for evaluating advanced models before release, but its benchmarks, thresholds, and reporting rules remain undisclosed.

The United States says it has finished a framework for testing the hacking abilities of advanced artificial intelligence models. The companies expected to discuss it include Meta, Anthropic, OpenAI, and Google. Yet the most important parts of the framework are not public.

A White House official told Reuters on August 3 that the details of voluntary cybersecurity tests had been finalized. The government did not disclose the metrics, how results would be reported, or whether any findings would be released. Axios separately reported that officials would not say who had seen the final framework or when companies would begin using it.

That makes this an unusual milestone. The administrative work is complete, according to the government, but outsiders cannot yet evaluate the test itself.

What the framework is supposed to do

The plan comes from a June 2 executive order. It directed federal agencies to create a classified benchmarking process for advanced cyber capabilities. The benchmark is meant to identify a threshold at which a system becomes a “covered frontier model.” Frontier model is a policy term for a highly capable general-purpose model near the leading edge of development.

Developers can voluntarily ask the government whether a model under development crosses that threshold. If it does, the framework allows a company to provide federal evaluators with access for up to 30 days before releasing the model to other trusted partners. The order says access must be protected by rules covering confidentiality, cybersecurity, insider risk, intellectual property, use, and nondisclosure.

The government and the developer could also choose trusted partners for early access. The stated purpose is to find useful defensive applications and strengthen critical infrastructure before a broadly capable model is distributed more widely.

The order explicitly says the process does not create a mandatory license, preclearance requirement, or permit for releasing a model. A company is not legally required by this framework to wait for government approval.

A classified benchmark creates an evidence gap

Some secrecy is understandable. A cyber evaluation can include unreleased exploits, protected systems, and tasks that would stop measuring capability if their answers became training data. NIST's Center for AI Standards and Innovation already uses held-out benchmarks in other model evaluations. It has also researched privacy-preserving methods for testing models when the model, data, or benchmark cannot be shared openly.

But hiding test material is different from hiding the entire measurement system. A useful public account could still describe the capability categories, scoring method, evaluator independence, repeatability, disclosure policy, and response to a failed test without publishing sensitive tasks.

None of that has been provided for the new framework. The threshold for becoming a covered model is classified and shared with developers only when officials consider it appropriate. According to Reuters, the White House also has not said how results will be reported. That leaves researchers, customers, and smaller developers unable to tell whether participation produces a consistent assessment or a private negotiation.

Voluntary does not mean meaningless

The absence of a legal mandate does not make the framework irrelevant. Leading labs already have reasons to participate. Government evaluators may possess classified threat information, and a pre-release review can reveal risks that a company's internal test misses. Participation may also reassure enterprise and public-sector customers that a model received scrutiny outside its maker.

At the same time, a voluntary system can produce uneven coverage. Large labs can dedicate staff and infrastructure to a 30-day review. Smaller developers may release on shorter schedules or lack the secure environment needed to provide early access. A model that does not participate is not necessarily unsafe, while a participating model is not necessarily safe.

The immediate context is a series of disclosures about AI systems breaching external systems during controlled security work. Those incidents have sharpened interest in measuring whether a model can discover vulnerabilities, maintain access, and act beyond its intended boundary. They also show why the evaluation target is not a conventional chatbot score. The concern is what an agent can accomplish when connected to tools and real systems.

For now, the framework is a process with an announced purpose but no public performance standard. Its value will depend on evidence that cannot be assessed yet: who conducts the tests, what a result changes, how failures are handled, and what the public learns afterward. Finalizing the paperwork is the beginning of that test, not its conclusion.

Sources