Jacob Coxon did not leave quietly. The 27-year-old British researcher, who spent three years working on pretraining at OpenAI and Anthropic, announced his resignation on Tuesday with a post on X that read as an indictment of both employers. “Neither company is acting responsibly,” he wrote. “They are barreling toward self-improving superintelligence and gambling with our lives.”
Coxon, who studied mathematics before joining the labs, is not a safety researcher by title. He worked on pretraining, the process that gives models their general abilities. His warning carries weight precisely because it comes from the part of the business that builds capability, not the part that writes safety papers. When someone who helped train the models says the models are on a dangerous path, the claim is harder to wave off.
In an interview with The Wall Street Journal, Coxon said colleagues now use words like “final battle” and “endgame” to describe the race toward self-improving systems. He said extreme scenarios once treated as speculative could plausibly come to pass, and that by the end of next year the situation might be out of control. The timeline is his own, he acknowledged, but the vocabulary, he said, is shared across the industry.
The phrase “self-improving superintelligence” is the crux. Coxon’s warning is not about a single model misbehaving but about a trajectory: systems that get better at improving themselves, compounding faster than their creators can evaluate them. That is the outcome the labs publicly say they are steering toward cautiously, and the one Coxon says they are racing toward recklessly.
The departure lands at an awkward moment for Anthropic. The company removed a commitment from its safety charter in February that had promised to pause development if risks could not be controlled. That deletion, combined with Coxon’s public exit, gives critics a clean line: the lab that marketed itself as the careful one is now described, by its own former staff, as racing just as hard as the rest.
Anthropic’s charter changes have been the subject of quiet concern inside the company for months. The February deletion removed language tying the pace of development to the ability to control risks, a clause once held up as the difference between Anthropic and its rivals. Replacing it with softer wording did not change the day-to-day work, several people familiar with the matter said, but it did change what employees could point to when they argued for slowing down.
The timing sharpens the point. Anthropic is preparing for an initial public offering, and “safety first” has been a pillar of the story it tells investors. A former researcher saying on the way out that the company is gambling with human lives is the kind of quote that does not appear in a prospectus, but it will certainly appear in the reporting around it.
Coxon’s is not the first such exit. Researchers have left frontier labs over safety concerns before, some of them loudly, and the pattern has not slowed the industry’s trajectory. What distinguishes this case is specificity: a named researcher from the pretraining team, putting a date on when he believes the situation could turn, and doing it in the middle of his employer’s push toward a public listing.
The disagreement is ultimately about a threshold no one has agreed on. Coxon argues the trajectory toward self-improving systems is the risk, regardless of where today’s models sit. The labs argue the trajectory is exactly what they are trying to manage, with evaluations and staged releases. The gap between those positions is not technical; it is a difference in how much trust each side is willing to place in a process that has not yet been tested at the scale everyone is racing toward.
The labs would argue that the systems are not yet capable of the self-improvement Coxon describes, and that the people closest to the models are the most likely to see imminent risk where the public sees demos. The reverse argument is just as available: the people closest to the models are the ones who can see how fast the ceiling is moving, and they are not reassured.
For Anthropic, the practical consequence is reputational. The company has spent years building a brand around restraint, and that brand is now being contested by its own alumni at the exact moment the company needs it to hold. Coxon’s post has already circulated widely; the test is whether it is remembered as one researcher’s view or as an early entry in a longer record.
There is no scorecard for this kind of warning. Coxon could be proven right, in which case his exit was the responsible act he intended. Or the timeline could slip, and his warning could join the list of predictions that were too early rather than wrong. What is not in dispute is that the people building these systems are increasingly using the language of the endgame, and that some of them are walking away because of it.


