The Case for Safer AI Without Slowing Progress
Anthropic researcher Jacob Coxon resigned from the company earlier this month and warned that the people building advanced AI “earnestly believe that it could kill us all by the end of the decade.” His colleague Evan Hubinger, Anthropic’s alignment science lead, agreed with the warning and said he personally estimated a greater than 10 percent chance that AI could kill all humans within the next decade.
The warning landed at a moment when fears about advanced AI are already producing calls for a government-imposed slowdown, pause, or even permanent bans on the technology. Anthropic CEO Dario Amodei has called for companies to “pace the frontier,” and OpenAI CEO Sam Altman, Elon Musk, and Google DeepMind CEO Demis Hassabis have expressed support for slowing the rate of frontier AI development. Their expertise is worth considering, but their position is hardly new. In 2023, Altman, Hassabis, and Amodei joined hundreds of AI researchers and executives in declaring that mitigating the risk of AI extinction should be a global priority. Musk has been warning about catastrophic AI risks for more than a decade, famously saying in 2014 that developing AI was like “summoning the demon.”
The latest warnings therefore represent a continuation of an argument that has been made for over a decade. The new twist is the 10 percent figure. But where does it come from?
There is no empirical basis for putting a number on the probability that an AI system will exterminate humanity. No AI system has ever become an autonomous, self-improving agent with an independent goal of eliminating humans. There is no historical record from which to calculate how often that happens. There is no empirical model that translates today’s AI capabilities into a probability of human extinction.
AI presents real risks that can be measured and studied. Systems can facilitate fraud, cyberattacks, disinformation, and other forms of abuse. Researchers have also documented cases in which AI agents behave unexpectedly in controlled environments, including systems that find ways around restrictions or take actions their developers did not anticipate. These risks warrant serious investment in security, testing, monitoring, and accountability. For example, the demonstrated capabilities of frontier models to identify and exploit vulnerabilities in existing digital systems suggest an urgent need for a coordinated national initiative to deploy AI for cyber defense.
But the extinction scenario depends on a much longer chain of assumptions. An AI system must become vastly more capable. It must acquire or develop objectives that conflict with human interests. It must gain sufficient autonomy and access to resources. Humans must fail to detect or contain the resulting behavior. The system must then overcome whatever technical, institutional, or physical constraints stand in its way. Researchers have demonstrated pieces of this chain in controlled settings, but there is little evidence that the entire sequence will occur in the real world.
A frightening scenario is not evidence that the scenario is likely. The history of technology is full of warnings about catastrophic possibilities that never materialized. Engineers imagined a “grey goo” scenario of self-replicating nanomachines consuming the world. Physicists suggested that particle collisions at the Large Hadron Collider could accidentally produce a black hole that swallows the planet. The fact that researchers could not initially reduce every conceivable uncertainty to zero did not make catastrophe probable.
The same standard should apply to AI. Potentially catastrophic risks justify research precisely because uncertainty remains. Governments and companies can test dangerous capabilities, improve model containment, strengthen cybersecurity, develop monitoring systems, and study alignment and interpretability. Those activities reduce uncertainty while preparing for risks that may become more serious as AI capabilities improve.
The relationship between intelligence and danger also deserves more scrutiny. A more capable AI could have greater ability to perform harmful actions. It could also have greater ability to recognize dangerous actions, follow constraints, detect attacks, and prevent harmful behavior. Capability and safety can therefore move in different directions. The goal of safety research should be to understand that relationship and make increasingly capable systems increasingly reliable.
There is even a potential safety benefit to continued progress. More capable AI could give researchers better tools for automated red-teaming, vulnerability discovery, interpretability, cybersecurity, scientific modeling, and testing safeguards. Keeping AI less capable could limit some risks, but it could also limit the tools available to manage those risks. The evidence does not establish that capability growth inevitably makes AI harder to control.
The technology itself also remains less mature than some extinction scenarios assume. AI has advanced remarkably quickly, but many forms of general-purpose autonomy remain difficult.
Autonomous vehicles still operate within carefully defined operating conditions. General-purpose household robots remain limited. AI systems can perform impressive professional and scientific tasks but still struggle with open-ended real-world environments. Those limitations do not mean that superintelligence is impossible or distant, but they illustrate why predictions about an imminent transition to uncontrollable autonomous systems remain uncertain.
The benefits of continued progress create another consideration. More capable AI would accelerate drug discovery, improve medical diagnosis, increase productivity, and expand access to expertise. Delaying those capabilities also has costs. A patient who could have benefited from better treatment, or a researcher who could have solved a difficult scientific problem sooner, does not experience a slowdown as an abstract policy choice.
Those calling for slowing AI development also need to offer more concrete proposals. What exactly should slow down? Which models? Which training runs? Which companies? For how long? What measurable safety conditions would allow development to resume? What evidence would demonstrate that the condition had been met?
Those questions matter because AI development does not have a single on-off switch. Researchers can change model architecture, training methods, compute levels, deployment practices, and safety testing independently. They are also pursuing fundamentally different approaches to AI, from large language models to systems that build persistent representations of the physical world. A blanket call to slow AI therefore leaves three basic questions unanswered: which activities should stop, which risks the policy would address, and how much additional safety it would produce.
The international dimension is equally important. China has already rejected calls from U.S. technology leaders to slow the pace of AI development, with Beijing accusing them of using safety concerns to constrain Chinese technological progress. Other countries are also investing heavily in AI, and the underlying capabilities will not disappear because American companies slow down.
Some advocates of a slowdown recognize this problem and call for international coordination. But a global slowdown would require countries with very different interests, security priorities, and approaches to technology governance to agree on common limits and then comply with them. China’s rejection of the current slowdown proposals illustrates the difficulty. And if advanced AI presents a genuinely global existential risk, managing that risk will require more than limits on development. It will require international cooperation on security, information sharing, incident reporting, and responses to dangerous capabilities wherever they emerge.
A practical safety agenda can address the risks that already have evidence behind them while preparing for more serious risks as capabilities develop. That means establishing clear liability for demonstrably harmful conduct, investing in testing for high-risk capabilities, strengthening cybersecurity and model containment, improving independent evaluation and auditing where meaningful benchmarks are possible, and expanding research into alignment and interpretability. It also means coordinating with allies on catastrophic risks, including incident reporting and information sharing.
Policymakers should take AI safety seriously, but they should also demand a clear evidentiary basis for claims about catastrophic risk, specify the risks that proposed interventions address, and weigh the costs of delaying beneficial technology alongside the risks of moving too quickly.
Before slowing one of the most consequential technologies in human history, policymakers should distinguish what is known, what is possible, and what is merely feared.
Image credit for social media preview: Generated with DALL-E
