Don’t Slow Down AI Development—Speed Up Benchmarking

Every news outlet has lately been seduced by the clarion call of P(doom) and the idea that rapid advances in AI could lead to catastrophic outcomes. Led by a mix of CEOs and activists, these calls have precipitated nothing short of a wave of deep public anxiety and fear about how AI will disrupt the future.

Yet, fear always misses the total picture and tends to obscure more than it clarifies. For all the recent hype regarding existential risk, the simple fact is that the United States must speed up both AI development and the checks and balances required to ensure that the technology produces more public goods and profit than pain. That shouldn’t mean halting development or trusting companies alone to build in safety standards absent public debate. Instead, it implies a need for more robust investment in AI research, including benchmarking, and the creation of new policy labs that study how to implement AI into critical national security workflows.

The United States should pair faster AI development with sustained investment in independent evaluation; this can be accomplished in several ways. Congress should expand basic safety research through the National Science Foundation (NSF) while enabling the National Institute of Standards and Technology (NIST) to coordinate public-private partnerships to develop common benchmarks. Universities and nonprofits, meanwhile, should host policy labs that test AI in high-risk national security workflows. Federal grants and foundation matching funds could support this work, alongside consideration of a token-use levy to finance public-interest research. These domestic measures should be coupled with AI-focused diplomacy that advances common evaluation standards globally and ensures that the AI arms race between the United States and China does not end in a security dilemma.

Abundance of Fearmongering

The debate over P(doom)—the probability of an AI-driven existential catastrophe—includes striking forecasts from the technology’s leading figures. In December 2024, Geoffrey Hinton, the “godfather of AI,” estimated a 10–20 percent probability of AI causing human extinction within 30 years. Elon Musk offered a similar estimate in 2024. In September 2025, Anthropic CEO Dario Amodei put the chance of things going “really, really badly” at 25 percent. Nonprofits focusing on AI safety have translated these concerns into demands for collective action. The Center for AI Safety’s 2023 statement, for example, signed by Amodei, OpenAI’s Sam Altman, and Google DeepMind’s Demis Hassabis, placed preventing AI-driven extinction alongside addressing pandemics and nuclear war as a global priority. In 2023, an open letter by a mix of AI experts called for a six-month pause in training systems more powerful than GPT-4. The Machine Intelligence Research Institute claimed that superintelligence developed soon using current methods would most likely cause human extinction, warranting a globally enforced moratorium.

While striking, these proclamations shouldn’t automatically be accepted as harbingers of doom. In one sense, the surge of fear could be clever market positioning, with companies acting like champions of safety for the simple act of regulating themselves. In another, it could be genuine fear and a concern for the common good. The CSIS Futures Lab itself identified key escalation risks in the AI model behaviors. While top engineers in the frontier AI labs might have much better insights regarding model behaviors than the watchers, doomsday scenarios nonetheless are not moving the debate. Circulating risk estimates and institutional positions warrant scrutiny. Their differing definitions, time horizons, and assumptions about future safeguards undermine their credibility. The lack of clear metrics of the tendency of AI model risks or agentic risks leaves citizens in the dark. This is why, regardless of the motive, both the public and private sectors in the United States should invest in independent AI evaluations.

Proposing safety measures is the right direction. In his latest essay “We Must Pace the Frontier,” Amodei proposes solutions operating at three levels. The first would involve embedding independent evaluators with employee-like access to training processes in frontier labs and the right to publish unfavorable findings. Second, laboratories in democracies would coordinate around common safety standards and capability-triggered checkpoints, supported by regulation and narrowly tailored antitrust waivers, where necessary. Third, governments would pursue verifiable international agreements, including with China, beginning with restrictions on dangerous uses and pre-deployment testing before attempting limits on recursive self-improvement, where AI helps develop more capable successors.

There is substantial agreement in the AI community on the need for human control alongside competing theories of how to preserve it. Altman has endorsed pacing, independent audits, and consistent federal standards, arguing that companies should begin without waiting for legislation or antitrust exemptions. Hassabis proposes a federally overseen frontier-AI standards body, while Microsoft sets out requirements for future models to accept human correction and shutdown. Mark Zuckerberg proposes board-level oversight of model releases and early government access to developing models.

Evaluation of AI Systems Is the Top Priority

AI is a general-purpose technology, and it should produce public goods while supporting the type of wealth creation and creative destruction that aligns capitalism with progress. Since the emergence of civilization, governments have struggled to keep up with the private sector in terms of regulating technology and commerce. The Code of Hammurabi, which is over 3,700 years old, had provisions for punishing third parties for building faulty houses. What would the Babylonian king say about AI?

AI benchmarking is the way ahead. It is central to keeping AI aligned with human values and exposing where machine logic drifts in unique use cases such as national security. If independent institutions can set AI evaluation standards, they can establish the foundation of trust in AI systems. This is also true for military use cases. With the lack of national security benchmarks in defense, the public remains largely in the dark about the specific potential risks of AI in military applications. For example, there are clearly missing benchmarking standards for understanding how AI model outputs align with the Laws of Armed Conflict. As such, U.S. foreign policy should also prioritize pushing international benchmarks as part of its AI diplomacy to prevent any security dilemma in the international system.

The federal government can help AI development speed up in a responsible manner by creating incentives for the type of safety research and benchmarking required to balance the rollout of increasingly powerful AI models. That should be done through public-private partnerships, similar to what NIST created through the Center for AI Standards and Innovation and the Trump administration’s call for X-Labs built around grand research challenges. The push to support deeper AI safety research should also see more basic research grants administrated through the NSF, which had its budget cut by over 50 percent in recent years, to more traditional universities and non-profits. These efforts could encourage more private sector support, such as matching grants with foundations and even creative mechanisms such as a Tobin tax on tokens to support basic research into AI safety standards and public use cases.

At their core, these efforts should build on the Anthropic CEO’s call for embedded third-party evaluators, common safety standards across democracies, and the establishment of venues to coordinate with authoritarian regimes to avoid an AI arms race. What is missing in that discussion are the ways and means. The United States needs to champion benchmarking and use-case studies around high-risk areas such as national security questions, large-scale cyberattacks, and bioterrorism. Such efforts must embrace a wider network than just internal researchers at frontier labs, which is why proper guardrails and benchmarking will be a key check and balance.

Conclusion

Responsible acceleration requires AI evaluation capacity to grow alongside the technology itself. Predictions of catastrophe offer limited guidance about how particular systems will behave in consequential national security decisions. However, independent benchmarking and sustained research can be critical tools, even panaceas, for identifying failure conditions, testing safeguards, and establishing a stronger basis for deployment decisions. Their value will depend on rigorous methods, meaningful access to models, and continued scrutiny as capabilities change. The federal government should fund this research, broaden participation beyond frontier laboratories, and make shared evaluation standards a priority for AI diplomacy. Public confidence should ultimately rest on evidence that institutions can challenge and update as AI advances.

The upcoming Trump-Xi summit offers an opportunity for progress on international governance of AI development. China’s spy chief recently warned the Chinese Communist Party regarding the political risks of advances in the AI domain and suggested that China should invest more in controlling its development. This illustrates why Trump and Xi will likely bring AI into their discussion, and neither side seems primed to show much appetite for slowing down AI progress. The better solution would be building a governing body that includes members from both powers that could coordinate the governance measures for good.

Benjamin Jensen is director of the Futures Lab and a senior fellow for the Defense and Security Department at CSIS. Yasir Atalan is a deputy director and data fellow in the Futures Lab at CSIS.

Image
Benjamin Jensen
Director, Futures Lab and Senior Fellow, Defense and Security Department
Image
Yasir Atalan
Deputy Director and Data Fellow, Futures Lab, Defense and Security Department