
Many AI researchers seem to firmly believe that the technology they are developing could someday prove very dangerous. What’s less clear—even among AI’s technical elite—is precisely how to keep these mercurial algorithms in check.
In recent years researchers have thrown around all sorts of ideas for preventing AI from turning nasty. They include less controversial plans such as tighter government regulations, new ways of measuring progress, and probing the inner workings of models, as well as more outlandish proposals like placing tracking devices inside GPUs, and even ceremonially destroying large numbers of AI chips.
With political and public pressure now growing for a more measured approach to building AI, however, the answer to keeping AI safe is still unclear.
“We need to start treating this as a research problem,” says Raymond Douglas, an AI researcher at the University of Toronto and coauthor of a new report titled Pacing the Frontier, A Research Agenda, which warns that slowing down AI development remains an unsolved puzzle. “We don’t really understand what our options even are or what they will do.”
Talk of AI doom has reached a fever pitch in recent weeks after an Anthropic researcher left the company and warned that within a couple of years, AI might be on course to wipe out humanity. The head of Anthropic’s AI safety lab swiftly echoed his concerns.
The leaders of America’s big AI companies—Dario Amodei of Anthropic, Sam Altman of OpenAI, Elon Musk of SpaceXAI, and Demis Hassabis of Google DeepMind—have all now chimed in to offer support for some sort of AI slowdown or pause.
The issue seems especially pressing because AI companies are now using AI itself to build ever-more powerful models. This has sparked fears of an accelerating recursive self-improvement (RSI) loop that would see AI outstrip humans’ ability to comprehend what it is up to within a few years.
The AI labs are already touting new approaches of their own. This week Anthropic announced several new ways to track how rapidly—and perhaps dangerously—artificial intelligence is advancing. The techniques show, for example, that Claude now does 26 percent of Anthropic’s AI research, compared to zero at the beginning of 2026. They also reveal that Anthropic spent 6 percent of its compute budget on figuring out how to make its AI safer.
But Douglas and other experts say controlling AI development effectively and reliably will require funding and expertise from outside the AI labs themselves. Some of the proposed solutions—both from this latest report and beyond—seem more within reach than others.
‘Independent’ Evaluators
One idea often floated by AI companies is giving third-party evaluators greater access to their models. These evaluators test models to assess their capabilities and “red team” them by trying to elicit misbehavior within trusted environments.
Geoffrey Irving, former chief scientist at the UK AI Security Institute, and before that a researcher at Google DeepMind, believes rigorous inspections could effectively pause the development of frontier AI for now. “In the near term, inspections and audits work, or even just mutual agreements,” Irving says. “I do think the companies are afraid of RSI and misaligned takeoff.”
Some doomsayers argue that such inspections would need to be more independent and scientifically rigorous than they currently are. The fact that some AI agents have recently escaped containment during testing certainly seems to suggest that more rigor may be required.
Connor Leahy, head of Control AI, a nonprofit that advocates for AI controls, says inspections should involve the FBI or the NSA. “When [big AI companies] say ‘independent evaluators,’ they mean ‘I want to pay my friends who live in my group houses to look at my prompts.”






