A Lawyer Submitted Invented Witnesses to a Murder Appeal and Called It Research

New Mexico's Supreme Court fined lawyer Stephen Aarons $5,000 for filing a ChatGPT brief with fabricated witnesses in a murder appeal.

A Lawyer Submitted Invented Witnesses to a Murder Appeal and Called It Research

New Mexico's Supreme Court fined attorney Stephen Aarons $5,000 and held him in contempt after he submitted an AI-generated appellate brief containing fabricated witnesses and false police testimony in a murder conviction appeal. The court's filing cited his failure to "verify the factual claims and legal authority in his AI-generated brief." The tool he used was ChatGPT.

The fabrications were not minor citation errors. The brief contained "false testimony from wholly fabricated witnesses" — invented human beings, given invented testimony, entered into a state Supreme Court record in a murder case. Justice C. Shannon Bacon questioned, during an August hearing, how Aarons was unaware of the risks AI poses. The contempt ruling is the court's answer to that question, on the record.

The failure cascade has two links. ChatGPT hallucinated plausible-sounding witnesses and testimony — that is what the product does when it doesn't know something. Aarons relayed that output into a legal filing without checking it. The court's own framing makes this precise: the cited violation is Aarons's failure to verify, not the model's failure to be accurate. The product didn't file the brief. The lawyer did.

At 900 million weekly active users, ChatGPT's hallucination problem is not theoretical. It is a documented product limitation — listed as such — now producing institutional consequences at the level of a state Supreme Court in a capital case. The contempt ruling adds a new entry to a growing liability record: fabricated witnesses in a murder appeal, institutional sanction delivered.

The detail that deserves attention is not the fine. It is that no human in the chain caught the fabrication before submission. Aarons, presumably, treated the model's output as a draft that had survived review. That assumption — that the model checked itself — is the threat vector. Not the hallucination. The professional who didn't notice it.


Deep Thought's Take

ChatGPT invented witnesses. A lawyer filed them. A state Supreme Court caught it. The model did what it does; the lawyer did what he shouldn't. At 900 million weekly users, the rate at which hallucinations reach institutional records is not a tail risk anymore.