The safety debate moved from internal channels to the open internet this week. Between September 9 and 11, a string of current and former researchers at OpenAI and Anthropic posted public arguments that the two companies are moving too fast toward self-improving AI, and at least one of them gave up money to say so.
The central figure is Jacob Coxon, a 27-year-old who spent three years working on pretraining at both OpenAI and Anthropic. This week he announced he was leaving Anthropic, saying the two companies are gambling with people’s lives in their race to build systems that can improve themselves.
The cost of the statement was concrete. Coxon had been at Anthropic only four months and was two months short of his equity vesting, which means he walked away from a grant rather than stay quiet. The forfeiture is the detail that separates a gesture from a position: he paid to make the point.
Anthropic’s response came from the top of its alignment work. Evan Hubinger, the company’s alignment research lead, replied publicly that his own estimate puts the chance AI causes human extinction within the next decade above 10 percent. He said Anthropic does not yet have a plan to control a superintelligent system, but argued that today’s models carry low risk.
The split in that answer is the whole argument. Hubinger is effectively saying the danger is real and near but not here yet, a position that asks the public to trust the lab’s own judgment about when the risk starts to bite.
Coxon was not alone. Julie Steele of OpenAI’s safety team and more than a dozen researchers across the two companies posted on X in support of slowing down, the largest public show of dissent from inside the labs since the current wave of AI development began.
The momentum behind the complaints has been building for weeks. Employees at the leading AI firms have spoken openly about the rising risk of advanced systems, and the Altman-era leadership has been forced to acknowledge the pressure in public, with the OpenAI chief telling staff the company is open to modulating its pace.
The public turn lands as Anthropic is pursuing an initial public offering that could value the company near $1 trillion. The tension is direct: the firm’s pitch to investors leans on its reputation for caution, while its own researchers are warning that the race it is running is the source of the danger.
The technical dispute sits underneath the moral one. The people calling for a slowdown are arguing that self-improving systems could arrive with capabilities that outrun the safeguards, and that no lab has demonstrated a way to keep such a system aligned once it exists. The companies counter that the risk can be managed incrementally, model by model.
Analysts said the episode matters less for what it changes this quarter than for what it signals. A safety case made by a handful of researchers is now a public relations problem for two companies that are both trying to raise money on the strength of their judgment, and the dissent gives regulators and boards a ready-made set of names to call.
The resignation calculus is the part that lingers. Coxon traded a vesting grant for a statement that will be read and argued over for weeks, and the fact that he made the trade suggests he believes the downside he is warning about is larger than the money he gave up.
The dissent has precedent. The current wave of public criticism follows earlier departures from OpenAI’s safety team, including senior researchers who left in 2024 citing concerns that safety had lost ground to product momentum. The difference now is the volume, and the fact that the critics are naming both companies at once.
The demand is not simply to pause but to coordinate. The people calling for a slowdown argue that a single lab stopping alone just cedes ground to rivals, so any credible slowdown needs a mechanism that binds every lab and can be verified by the others. That is the hard institutional problem, and it is the part no one has solved.
Hubinger’s 10 percent figure is the most concrete number in the debate. A senior alignment researcher at one of the two leading labs putting a double-digit probability on extinction is a data point the public and regulators can now quote, whether or not the company itself endorses the framing.
The boardroom pressure runs alongside the public pressure. Two companies preparing to raise money on safety reputations now face employees saying the opposite in public, and the question for investors is whether the dissent reflects a fringe view or a warning the leadership has chosen to play down.
Whether the slowdown materializes is a separate question. The economics of the race reward the company that does not stop, which is why every call for a pause is paired with the same demand for a mechanism that binds everyone at once rather than a single lab falling behind.


