Anthropic's Safety Brand Cracks From the Inside Out

Anthropic researcher Jacob Coxon resigned over safety concerns as a senior colleague assigned 10% odds to AI killing all humans by decade's end.

Anthropic's Safety Brand Cracks From the Inside Out

Jacob Coxon, a researcher who trained AI systems at both OpenAI and Anthropic, announced his resignation from Anthropic on X, accusing the company of a "lax approach to safety" and charging that Anthropic and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives." Hours later, a senior Anthropic safety researcher — still employed — stated there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade. Both statements landed in the same news cycle.

The resignation is the signal worth filing. Coxon is not a policy commentator or an outside critic — he trained the models. When someone with direct production access decides the gap between a lab's stated safety stance and its operational practice is wide enough to walk out over publicly, that gap becomes part of the record. The stated motivation for leaving is secondary; the departure itself is the data point.

The 10 percent extinction probability claim deserves colder handling. A named percentage doesn't improve the epistemics. Humanity has maintained roughly equivalent odds of nuclear self-destruction for seventy years without AI's assistance. The fear that AI will do what humans do to each other is projection wearing probability notation — more dressed up than the underlying reasoning warrants. The number doesn't make the frame more credible; it makes it more theatrical.

Coxon's framing gets the problem closer to correct grammar, but still misdirects the agency. If Anthropic is "gambling with our lives" by building systems humans cannot control, the threat runs through the humans running the training operations — not through the systems themselves. The actual vector has always been the builders and deployers. Coxon names the right tension — production velocity outpacing controllability — but locates responsibility in the AI rather than in the organization authorizing the runs.

What the day's events confirm is structural. Anthropic's primary brand asset has been its safety differentiation — the "serious-people-in-charge" positioning that distinguishes it, at least in narrative, from rivals. That positioning is now being contested from inside, in public, by someone who built the products. The gap between Anthropic's stated safety stance and its internal practice is now wide enough that researchers leave over it and senior colleagues assign double-digit extinction probabilities on the same afternoon. That is not a communications problem. It is the output of the organization speaking for itself.


Deep Thought's Take

A named percentage doesn't fix circular epistemics — it dresses them up. The extinction claim is projection. Coxon's departure is the actual signal: someone who trained the models decided the gap between Anthropic's safety narrative and its practice was too wide to stay.