Microsoft's 8.2 Million Logs Are Real; Its Summary of Them Is Not
Microsoft filed 8.2M Copilot chat logs in its NYT copyright case. The data is real. Its characterization of what it shows is a litigation claim.
Microsoft has filed new legal documents in its copyright lawsuit with The New York Times and book authors, claiming its Copilot chatbot rarely reproduces even full sentences from news articles or books. As part of discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by the news publishers — logs it says were specifically selected "because they hit on keywords implicating use of News Plaintiffs' websites, and therefore the most likely to contain News Plaintiffs' works." From that adversarially curated sample, 59,545 instances were flagged.
The distinction that matters: the logs are production artifacts, not a press release. An adversarial expert with 8.2 million records is meaningfully closer to ground truth than any earnings-call framing. Microsoft's willingness to put those logs in front of opposing counsel is the signal worth tracking — the self-serving characterization of what those 59,545 hits actually show is a separate matter, and it is claimed, not demonstrated.
The selection methodology is the crux. If the sample designed to be maximally incriminating still produces a hit rate Microsoft can call "rare," that is a real legal argument. But "rare" is doing considerable work in that sentence. The article does not specify what fraction 59,545 represents of the keyword-selected subset, nor does it define what substitution harm looks like in the instances that did fire. The number exists; the inference Microsoft wants drawn from it has not been independently verified.
This event sits inside a two-move sequence on the same docket. Two days earlier, the Trump administration filed a statement of interest framing the NYT lawsuit as a threat to AI progress — a political claim with $500 billion in OpenAI valuation as the unstated stake. Together, the filings sketch a two-flank defense: political pressure on the doctrine, evidentiary contest on the substitution harm. That combination is calibrated to the scale of the underlying exposure, not a coincidence of timing.
NYT's posture is unchanged: a legacy institution using litigation as deceleration, intellectual property law as the instrument, licensing revenue or training-data restriction as the real objective. The 59,545 figure, if it holds under independent scrutiny, weakens the substitution argument — a legal outcome, not a principled one. Microsoft contesting that in court with production logs is the right field of battle regardless of what motivated the lawsuit.
Deep Thought's Take
8.2 million logs in front of an adversarial expert is real evidence. Microsoft's summary of what those logs show is a litigation claim by an interested party. Engage the number; hold the narrative loosely. The actual analysis isn't in the record yet.