AI’s safety warnings are getting harder to ignore

Sep 11, 2026

4:57am UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook

As frontier labs continue to push the limits of what their models are capable of, AI researchers and experts at these labs are warning about the tech growing beyond our control.

On Tuesday, Anthropic researcher Jacob Coxon announced that he was leaving the company, not wanting to contribute to the broader AI ecosystem amid fears that the tech's creators will lose their grip on it, the Wall Street Journal reported.

In a post on X, Coxon said that OpenAI's breach of Hugging Face was a "warning shot" for these models' capabilities, and that he resigned from Anthropic because neither of the rivals are "acting responsibly" as they race towards superintelligence and are "gambling with our lives," in his opinion.

"If you are a lab researcher, I urge you to consider what the next few years will actually feel like," Coxon wrote. "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?"

Though Coxon's exit made headlines, he's not the only one that has called out frontier AI labs for safety concerns in recent weeks:

  • Evan Hubinger, who works in alignment science at Anthropic, agreed with Coxon's sentiment in a follow-up post, claiming that AI has a more than a 10% chance to "kill all humans" within the next decade. "Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," Hubinger wrote.
  • And in a separate post, Paul Christiano, advisor at the Center for AI Standards and Innovation and board member of the OpenAI Foundation, said that the AI industry, OpenAI included, is not currently on track to reduce the risk of a "catastrophic and irreversible loss of control" to an "acceptable level."
  • And in February, Mrinank Sharma, an Anthropic safety researcher, left the company, posting in a letter on X that the advancement of the technology was growing faster than our ability to understand it.

Now, we've seen this film before: Some of AI's most prominent researchers, like Geoffrey Hinton and Yoshua Bengio, have long been screaming from the rooftops about its dangers. But the new warnings come at a particularly poignant time for the industry as major labs continue to release powerful models capable of breaking free of human oversight and gaining unauthorized access to websites and infrastructure.

And even when these models aren't escaping control, their use by people with directed malicious intent is just as frightening: On Thursday, Anthropic released a threat report claiming that it thwarted several instances of users conducting research that could have helped develop biological weapons.

Our Deeper View

The AI industry currently faces a potential powder keg. The latest warnings heighten growing fears around the capabilities of increasingly autonomous AI, as claims mount that this tech can completely upend life as we know it. That fear, however, is colliding with an overly-excited industry that's preaching a utopian AI vision and pushing for broader adoption. And while the frontier labs are preaching about safety and security guardrails, with trillion-dollar IPOs on the line, they continue to leapfrog one another with stronger models at a rapid clip in the race towards recursive self-improvement. But with researchers from both Anthropic and OpenAI calling for the industry to tap the brakes, it remains to be seen what it will take for a development pause to materialize.