Search This Blog

Sunday, September 13, 2026

Warning!


Anthropic researcher Jacob Coxon resigned this week and left behind a blunt accusation on X: his former employers—Anthropic and, before that, OpenAI—were “gambling with our lives.”

Then, in a move no public relations team would have signed off on, Evan Hubinger, an alignment science lead at Anthropic, wrote that he agreed with Coxon and that he put the chance of artificial intelligence killing all humans within the next decade at greater than 10 percent. All this came from a senior researcher at a company that has been racing to build the very systems he fears.

In an interview with Wired, Coxon said incidents such as July’s Hugging Face breach—in which OpenAI models that were undergoing a cybersecurity test circumvented controls that were meant to isolate them from the Internet and compromised parts of the AI start-up Hugging Face’s systems—helped spur him to speak out, though he stressed that the problem was broader. (Coxon, Hubinger, Anthropic and OpenAI did not respond to requests for comment from Scientific American.)

Coxon’s and Hubinger’s fears, shared by plenty of other AI researchers, center on alignment, the difficult problem of matching an AI model’s behavior and apparent objectives to what humans want. The concerns still sound like science fiction: loss of control, recursive self-improvement and superintelligence—the possibility that AI models slip beyond their bounds, update their inner workings to gain power and eventually become too capable for humans to rein in.

Bruce Mehlman: