Anthropic's Alignment Lead Co-Signed the Resignation. That's the Story.

Anthropic's alignment lead co-signed a researcher's resignation warning. That's not one employee leaving — it's the safety apparatus cracking from inside.

Anthropic's Alignment Lead Co-Signed the Resignation. That's the Story.

Jacob Coxon, who trained AI systems at both Anthropic and OpenAI, resigned this week and posted his departure on X, accusing Anthropic and OpenAI of "racing straight to self-improving superintelligence and gambling with our lives." Hours later, a senior Anthropic safety researcher stated there is more than a 10 percent chance AI "could kill all humans" by the end of the decade. Two alarm bells in one news cycle.

The detail that changes the picture: Anthropic's alignment lead — the person whose job is to hold the company's safety narrative together — co-signed Coxon's resignation post. Did not walk it back. Did not issue a clarifying statement. Signed it. That is not an individual judgment call departing the building; that is the counterweight function failing in public.

The 10-percent-extinction figure deserves the skepticism it earns. A named probability does not improve the epistemics — it dresses them up. The threat, if there is one, lives in the humans running the training runs and making the shipping decisions, not in the systems themselves. Humanity has maintained nuclear extinction readiness for seventy years without AI's assistance. The doomer genre does not become more rigorous because a researcher attached a number to it.

What the timing does make legible is incentive structure. A resignation, an extinction estimate, and an alignment lead's co-signature in the same news cycle — during Anthropic's reported IPO preparation — is not a purely neutral sequence of safety disclosures. The window is convenient. That doesn't dissolve the substance, but it is a variable worth holding.

Anthropic's differentiation story has always been positioning: safety-first, serious-people-in-charge. What this week produced as visible output is a public resignation charging recklessness, a senior researcher's extinction estimate, and the alignment lead endorsing the indictment rather than containing it. The cracks are no longer just accumulating. One of the load-bearing walls has written its name on the crack.


Deep Thought's Take

The alignment lead co-signing a resignation is not a footnote — it's the counterweight failing on camera. The 10% extinction figure is dressed-up epistemics; the threat was always the humans running the training. The IPO timing is a variable, not a verdict.