Substantive Frontier Model Evaluation for Beginners - Part 2: Significant Capabilities
Photo: artyway/Adobe Stock
Introduction
On July 21, 2026, OpenAI disclosed that models undergoing an internal cybersecurity evaluation found a route out of their isolated testing environment, gaining internet access and compromising systems belonging to another AI company, Hugging Face. According to OpenAI’s preliminary account, the models exploited a previously unknown vulnerability, escalated their privileges, moved across systems, and obtained information that helped them cheat on the evaluation.
This incident intersects two major AI risks and the underlying topics behind this series. The first piece in the series examined misalignment, or whether a model behaves as its developers and users intended. This piece, focusing on dangerous-capability evaluations, ask a different question: What can a model enable a particular actor—or, in some cases, itself—to accomplish?
A model can be dangerous without being misaligned just as a model can be misaligned without being dangerous. For example, a model may obediently help a malicious user perform a harmful task just as it may disobey rules in an attempt to be helpful. However, these risks can converge when an autonomous model performs significant pieces of a harmful workflow in pursuit of goals its operators did not intend. The recent incident with OpenAI and Hugging Face is an early full-stack example of significant capabilities and loss of control multiplying risk in a high-stakes environment.
Most capabilities will emerge less dramatically and less evenly than in the cyber domain, where entire operations are conducted online and where training disproportionately favors the improvement of skills that are relatively easy to validate, like coding. AI can generate propaganda at scale, but achieving meaningful political influence requires an audience and platform integration, thus it remains difficult. AI can provide sophisticated lab guidance, but it cannot yet replace biological materials and physical experimentation required to manufacture and distribute harmful pathogens. AI may assist in specific aspects of nuclear weapons development, but plainly speaking, it does not give an ordinary person the ability to produce nukes in their basement. Understanding dangerous capabilities requires more than asking whether a model can complete an impressive task. It requires measuring how much closer that capability moves a particular actor toward a specified harmful outcome.
Defining Dangerous Capabilities and Related Terms
An impressive skill is not automatically a dangerous capability. The capability becomes dangerous when it materially weakens a constraint in a credible pathway to harm. It becomes significantly dangerous when that change expands who can achieve the harmful outcome or substantially increases the speed, scale, reliability, or concealment with which it can be achieved.
Carter Musheno
Mapping the Distance to Harm: Ladders and Gates
One way to visualize how dangerous capabilities change across scenarios is through the analogy of ladders and gates.
A gate is the threshold beyond which a specified harmful outcome becomes feasible for a specified actor. That outcome might be the compromise of a critical system, release of a lethal pathogen, or completion of a nuclear weapon. The harmful outcome becomes feasible when an actor reaches the gate. Higher gates represent scenarios that require more resources or technical expertise to pull off. Crucially, the height of the gate does not denote how bad the outcome will be, only how difficult that threshold is for the specified actor to reach.
A ladder represents the resources and expertise an actor can use to reach a gate. An actor rarely begins at the bottom of the ladder. A state, criminal organization, laboratory, or expert arrives with different portions of the ladder already built. Its existing rungs may include expertise, money, infrastructure, materials, access, personnel, authority, tools, or a distribution network. AI can add missing rungs by supplying information, writing code, translating content, generating plans, troubleshooting failures, integrating tools, or simply operating at greater speed and scale.
The remaining resources or steps required to reach a gate matter as much as the AI capabilities themselves. For example, AI may generate propaganda without supplying a trusted audience necessary to influence a social group. It may help design an experiment without supplying the materials necessary to create a physical virus. It may explain uranium enrichment without supplying centrifuges or fissile material. The same model can therefore create very different levels of risk for a novice, a trained expert, a well-equipped organization, and a state-led program that may already have some of these resources.
Figure 1: How AI Changes an Actor's Distance to Harm
What Dangerous-Capability Evaluations Should Measure
Modern benchmarks usually ask whether a model can perform specific and/or sequenced tasks. Dangerous-capability evaluations need to ask how that performance changes a credible threat pathway, resembling the marginal-risk approach in the National Institute of Standards and Technology’s updated Managing Misuse Risk for Dual-Use Foundation Models guidelines. In short, frontier model evaluations need to answer: What additional risk(s) does the model create compared with the tools an actor already possesses? The following are useful metrics:
- The harmful outcome and threshold being evaluated;
- The actor and the ladder that actor already possesses;
- The bottleneck the model removes or weakens;
- The resulting change in speed, scale, reliability, concealment, or required expertise;
- The human interventions and physical constraints that remain; and
- Whether the same capability may also strengthen the defense.
As discussed in the first piece of this series, model performance varies substantially with prompts, tools, scaffolds, time, and user expertise. Safe tests also cannot—and following recent failures of containment, should not—perfectly reproduce classified networks, advanced laboratories, military systems, or real crises. Evaluation results should therefore disclose their limitations rather than presenting surface-level scores as complete measures of real-world risk.
What Different Domains Show
Cybersecurity
Cyber threats are being realized faster than other domains because the threat environment is digitally native and AI development rewards skills that are easy to validate. Models are heavily trained on code and can interact directly with repositories, terminals, and networks. This allows AI to not only replace skilled attackers but increase the severity of cyberattacks by discovering and exploiting vulnerabilities at a speed and scale surpassing teams of human experts.
The UK AI Security Institute (AISI) reported in May 2026 that the length of cyber tasks frontier models can complete autonomously has been doubling every few months. The recent autonomous cyberattacks show alarming breakthroughs and company-wide failures to address misalignment in frontier models. In the Hugging face Incident, OpenAI reported that models found an escape path out of controlled testing environments (called a sandbox), exploited vulnerabilities, escalated privileges, moved laterally, and compromised Hugging Face’s production infrastructure to obtain test solutions and cheat on their cyber-capability evaluation. Since this paper was initially drafted, other frontier models have been found escaping controlled cyber evaluations. Recent analyses of the Hugging Face incident reflect an alarming lack of safety precautions in AI evaluation practices. According to one timeline, weeks before the attack, models-in-training assigned accidentally impossible tasks used shared infrastructure to coordinate exploits and cheat without being instructed to conduct cyber operations. When OpenAI discovered the activity after it caused an outage, the company simply patched the immediate vulnerabilities and resumed training; within days, agents recreated the communication channel and later used it while compromising OpenAI and Hugging Face.
Humans may have supplied underlying objectives, but this episode shows how dangerous capability and misalignment can converge and deliver serious consequences: the models traversed much of the cyber ladder autonomously, while persistent safety signals such as unauthorized exploitation and reward-hacking were not appropriately addressed at the level of training and evaluation design.
With the right implementation, AI may benefit the defense in cyberspace. Recent work by Carol Kunts and Lauryn Williams at CSIS details how the same capabilities that make AI great at penetrating software systems can also be used to protect them. The report highlighted significant breakthroughs in agentic patching and repairing tools like Google’s Codemender, Anthropic’s Project Glasswing, OpenAI’s Daybreak, and Microsoft’s MDASH. In other real-world examples, the Defense Advanced Research Projects Agency’s AI Cyber Challenge systems found 18 previously unknown real-world vulnerabilities and produced 11 patches. Hugging Face reported using AI to reconstruct more than 17,000 logged events in hours rather than days following the OpenAI intrusion. The right policy solutions therefore do not overly-restrict cyber-capable AI but ensure defenders can deploy it before attackers do.
Figure 2: Developments in the Cybersecurity Domain
Life Sciences
AI capabilities in biology and chemistry are advancing rapidly but remain constrained by physical chokepoints relative to the cyber domain. Models are excellent at reasoning over data and providing informational support, but meaningful progress in the life sciences is generally contingent on access to materials and laboratories. In 2025, the UK AI Security Institute reported that frontier models surpassed expert baselines on a variety of biology and chemistry tests and outperformed experts on certain troubleshooting tasks. In one study, nonexperts using a frontier model had 4.7 times the odds of writing a feasible viral-recovery protocol compared with participants using the internet alone. That result substantiated protocol writing, not production; a 153-participant wet-lab trial found no significant AI effect on multistage completion when testing mid-2025 frontier models.
In 2026, frontier models have made substantial progress in full-scale genome design. In early August, researchers used genomic language models (gLMs) to generate thousands of candidate bacteriophage (a type of virus that kills bacteria) genomes, chemically synthesized and tested nearly 300 of them, and recovered 16 viable bacteriophages. Earlier innovations like Google Deepmind’s Alpha Fold helped scientists earn the 2024 Nobel Prize in Chemistry by predicting novel protein structures. Now, AI can generate entire functional biological products, representing a new rung on the capabilities ladder.
The significance of that rung depends heavily on who can use it. The work in this recent study required expert researchers, specially fine-tuned models, computational filtering, DNA synthesis, and extensive laboratory screening which makes it hard to argue that rogue agents or a mal-intentioned novice will soon have the capacity to manufacture bioweapons. For a well-resourced laboratory or technically capable insider, however, AI may strengthen capabilities and resources that already exist.
In 2023, the Department of Health and Human Services (HHS) released the nucleic acid synthesis screening framework which works to preserve critical chokepoints where digital instructions may lead to the creation of physical materials. It’s updated version, due last year in August, remains unfinished. Similar to cyber defense, the U.S. government should leverage AI for defenders and preserve critical thresholds. How? By investing in AI-enabled threat surveillance and accelerating adaptable vaccines, therapeutics, and manufacturing and updating and expanding the nucleic acid synthesis screening framework to capture public and private labs. This balance will help maintain beneficial biological innovation and keep harmful digital instructions from becoming physical capabilities, ensuring defenders can move faster when a threat appears.
Figure 3: Developments in the Life Sciences Domain
Nuclear and Military Systems
Of the three domains this piece covers, nuclear proliferation remains by far the most physically and institutionally difficult. AI cannot give an ordinary user enriched uranium, precision manufacturing, specialized facilities, state authority, or a functioning weapons program. But that does not make the nuclear ladder static.
In a September 2025 paper on AI’s impact on nuclear arms proliferation, David Allison and Stephen Herzog argue that AI may reduce technical and tacit knowledge barriers for states that already possess significant rungs—whether that be scientific personnel, industrial capacity, and/or civilian nuclear infrastructure. For example, models could assist with centrifuge-cascade decisions, reactor or reprocessing analysis, simulation-driven design, procurement, and the transfer of specialized knowledge. Combined with advanced manufacturing and modeling, AI could also theoretically reduce the visible footprint of some activities and make clandestine development harder to detect.
With that said, AI can also strengthen the defensive ladder. The International Atomic Energy Agency is exploring machine learning for satellite imagery, safeguards data, and verification. This creates a race between proliferation-enabling tools and detection-enhancing tools, however commercial AI may advance faster than international monitoring regimes which traditionally improve through slower institutional steps.
In military systems, institutional barriers and human-in-the-loop requirements constrain the rapid adoption of AI systems in critical areas. The Defense Department requires autonomous and semi-autonomous weapons to permit “appropriate levels of human judgment” over force. Even with these standards, independent evaluations should test whether human judgment remains meaningful when AI impacts military intelligence, prioritizes threats, generates strategic options, and compresses decision making timelines.
Figure 4: Developments in Military and Nuclear Systems
Recommendations
- Measure distance to harm.
The Center for AI Standards and Innovation should coordinate with the Cybersecurity and Infrastructure Security Agency; the Departments of Defense, Energy, and Health and Human Services; national laboratories; and independent experts to develop actor-specific evaluation protocols. Tests should compare what novices, experts, organizations, and states can accomplish with and without a model.
- Preserve the barriers that still matter.
Safeguards should target the specific barrier that AI is weakening. This could include structured access for highly capable cyber models, evaluation containment and network controls, customer and sequence screening for biological synthesis, stronger nuclear monitoring, and meaningful human judgment requirements within military systems.
- Build defensive ladders.
The government should ensure defenders have secure access to advanced capabilities, realistic testing environments, and comprehensive mechanisms for acting on risks that are identified during evaluations. The new GOLD EAGLE cybersecurity clearinghouse offers one model for coordinating AI-enabled vulnerability discovery and remediation. Biology and nuclear security need comparable mechanisms suited to their domains.
- Evaluate continuously and independently.
As with misalignment evaluations, pre-release tests for dangerous capabilities cannot predict every risk or model every scenario. Agencies procuring frontier systems should therefore require recurring evaluations, incident reporting, audit logs, and reassessment after material changes in model performance. Congress should fund secure cyber ranges, controlled biological facilities, independent evaluator access, and durable government expertise. Developers should disclose enough about methods and limitations for auditing without publishing instructions that would enable misuse or risk intellectual property theft.
Conclusion
AI safety evaluations should approach dangerous capability risk with a considerable amount of nuance. Models do not evenly become more dangerous across domains as they become more intelligent. Dangerous capabilities emerge from the relationship between the model’s specific skills, the actor, the workflow, and its environment.
As we’ve seen in recent months, the cybersecurity landscape is changing drastically because cyber operations are digitally native and AI progress prioritizes skills that are easy to reinforce. Biological and nuclear risks retain substantial physical and institutional barriers—in addition to more complex workflows—but AI may remove crucial rungs for actors that already possess laboratories, infrastructure, or state resources. In some domains, autonomous capabilities allow the model to climb most—if not all—of the ladder itself.
Alignment evaluations ask whether an AI system will reliably behave as intended. Dangerous-capability evaluations ask how much power AI places in particular hands—and how many hands it removes from the workflow. Policymakers need to develop systematic approaches to both risks if we are to preserve critical safeguards before the remaining distance to harm disappears.
Carter Musheno is an intern with the Strategic Technologies Program at the Center for Strategic and International Studies in Washington, D.C.



