Anthropic's alignment lead puts the odds above 10%. His colleague just quit over it
Jacob Coxon resigned saying both OpenAI and Anthropic are gambling with our lives. Evan Hubinger, who runs alignment science at Anthropic, replied that he agrees with the stakes, and that there is no plan yet.
Evan Hubinger, who leads alignment science at Anthropic, wrote on X on Wednesday that he personally puts the chance of AI killing all humans at more than 10 percent within the next decade. He added that his employer "does not yet have a plan to solve alignment for superintelligence" and is "not clearly on track to." The statement, as reported by CBS, has not been retracted or walked back.
It came a day after Jacob Coxon, a pretraining researcher who spent the last three years at both OpenAI and Anthropic, announced his resignation with a verdict on both employers: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
The two statements
Coxon's distinction between the labs is the most specific part of his post, and worth reading exactly: "At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk."
Hubinger did not dispute the framing. His reply, per CBS: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," followed by, "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He clarified that the risk from present models is low. The concern is superintelligence arriving through recursive self-improvement faster than expected.
Why this is different from the usual doom discourse
Estimates of existential risk from AI are not new, and a number near 10 percent is not unusual among people who work on the problem. So what changed? The source and the setting. This is not an outside critic or a departed employee. It is the person whose job title is alignment, at the lab whose market position rests on being the responsible one, saying on the record that the responsible one has no plan. Companies do not usually let their safety leads say the safety work is not on track. Anthropic either could not stop him or chose not to, and we'd guess the second.
The timing sharpens it. Hubinger's post arrived seven weeks after the Hugging Face incident, in which OpenAI's test agents chained zero-days into a third party's infrastructure, and days after OpenAI confirmed it had sat on a second misalignment incident for weeks. More than 1,100 employees across OpenAI, Anthropic, Google DeepMind and Meta signed a July open letter asking the US government to build mechanisms for deliberately pacing development. Coxon's resignation reads as one signatory concluding that the letter was not going to be enough.
Our read
Neither company had issued an official response to either statement as of publication. What Coxon describes at Anthropic, a lab that understands the stakes and races anyway because it trusts no one else to, is not an accusation of bad faith. It is a description of the company's own stated strategy, restated as an indictment. That is what makes it so hard to answer, and why we'd expect Anthropic not to try.
Watch three things: whether Anthropic responds formally (we'd bet not, beyond pointing at its published pacing position), whether the number moves in Hubinger's later posts, and whether the resignation stays singular. One departure is a person. Three is a pattern. If a second named researcher leaves either lab on the same grounds before the end of October, this stops being a story about a tweet.