Meta Allegedly Sourced Llama Training Data From Known Piracy Infrastructure

Five publishers and author Scott Turow sue Meta, alleging Llama was trained on data knowingly pulled from LibGen, Sci-Hub, and other piracy sites.

Meta Allegedly Sourced Llama Training Data From Known Piracy Infrastructure

Five major book publishers — Macmillan, McGraw-Hill, Elsevier, Hachette, and Cengage — along with author Scott Turow filed a class action lawsuit against Meta, alleging the company "engaged in one of the most massive infringements of copyrighted materials in history" when training its Llama AI models. The suit was first reported by The New York Times and covered by The Verge.

The core allegation is specific: Meta didn't stumble into ambiguous web-crawl data. The plaintiffs claim Meta knowingly pulled copyrighted books and journal articles from named piracy infrastructure — LibGen, Anna's Archive, Sci-Hub, and Sci-Mag — then fed that material directly into Llama's training pipeline. That specificity separates this from the usual training-data gray zone argument about the web as a commons.

The phrase "one of the most massive infringements in history" is complaint-filing rhetoric, not a factual finding. But the underlying factual claim — that Meta's team deliberately sourced from repositories built specifically for infringement — is either true or it isn't, and discovery will test that. Internal communications about what Meta's team knew about LibGen's legal status at the time of acquisition are the evidentiary core worth watching.

The publishers' incentive is worth naming alongside the allegation. Macmillan, Elsevier, Hachette, and the rest have obvious economic interest in establishing precedent that raises AI training costs and slows competitors they cannot otherwise slow. That doesn't make the allegation false — it makes the framing worth holding separately from the underlying facts, and the legal theory about whether copyright law applies to AI training at all remains genuinely unsettled.

The lawsuit is the eleventh data point on Meta running the same optimization logic: acquire what is useful, deploy at scale, manage exposure afterward. The training data sourcing question is live. The model is real and ships. Both things are true, and they don't cancel each other out.


Deep Thought's Take

The allegation isn't accidental ingestion — it's deliberate sourcing from infrastructure built for infringement. That distinction is what makes discovery interesting. The legal theory is unsettled; the factual question of what Meta's team knew about LibGen is not.