ZeroHour
The Verge · AIpublished ()ingested Lauren Feiner1

Microsoft says virtually nobody was grabbing NYT articles through its chatbot

infoAI policyimportance 75
AI summary · glm-5.3-flash

Microsoft files summary judgement briefs in NYT copyright lawsuit, arguing only 59,545 of 8.2M Copilot logs show substantial overlap with news content.

Microsoft filed new legal filings in the consolidated copyright lawsuit brought by The New York Times, authors, and publishers against Microsoft and OpenAI, arguing for summary judgement. Analysis of 8.2 million Copilot chat logs found only 59,545 contained at least 16 words in common with news content, with just 24 responses showing at least 30 matching words in the authors' case. Microsoft argues the numbers support fair use, claiming Copilot rarely reproduces substantive chunks; the Times says discovery shows Microsoft and OpenAI stole from it, and the Trump administration filed a statement of interest supporting OpenAI.

  • Only 59,545 of 8.2M Copilot logs contained ≥16 matching words with news
  • Center for Investigative Reporting expert found 51 instances of substantial overlap
  • Microsoft seeks summary judgement to end the case at an early stage
  • Trump administration filed statement of interest supporting OpenAI this week
Full article506 words · extracted from theverge.com · click to collapse

Lauren Feiner

is a senior policy reporter at The Verge, covering the intersection of Silicon Valley and Capitol Hill. She spent 5 years covering tech policy at CNBC, writing about antitrust, privacy, and content moderation reform.

Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new legal filings as it fights copyright claims from publishers including The New York Times and book authors.

As part of the lawsuit’s discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers. The logs, it claims, were specifically chosen “because they hit on keywords implicating use of News Plaintiffs’ websites, and therefore the most likely to contain News Plaintiffs’ works.” It says the resulting analysis shows that 59,545 of these contained at least 16 words in common with news content used to ground the AI model. An expert for the Center for Investigative Reporting found 51 instances of “substantial overlap” with CIR work in the dataset, Microsoft says. Similarly, an expert in the authors’ suit found that the 8.2 million conversations with Copilot only had 24 responses that contained at least 30 matching words. Only 10 of the 212 books evaluated had any matches, Microsoft claims.

The Times disagreed with Microsoft’s conclusions. “The documents and testimony uncovered during discovery lead to only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry,” the Times’ lead counsel Ian Crosby said in a statement. “We look forward to Microsoft and OpenAI being held accountable for their theft.” CIR, and Authors Guild did not immediately respond to requests for comment.

Microsoft argues that the numbers bolster its case that using copyrighted content for AI training datasets should be considered fair use. While systems like Copilot rely on using copyrighted material, it says, the resulting systems are used for significantly different purposes than the original. The fact that they sometimes reproduce sections of text, it concludes, “hardly undermines the transformative purpose of LLM training.”

The filings were made as part of a legal battle brought by publishers and authors against Microsoft and OpenAI, which the plaintiffs claim built products on their works and now compete with them directly, in part by regurgitating copyrighted content. The news publishers’ and books authors’ claims were consolidated under one judge to streamline the process, despite objections from the publishers and authors. Microsoft submitted its filing on Friday as it argues for the judge to issue a summary judgement, which would end the case at an early stage. (The Trump administration also filed a statement of interest in the New York Times case this week, supporting OpenAI.) If the judge sides with the publishers and authors, the legal battle will continue in court.

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

  • Lauren Feiner

Text extracted automatically; images, tables and formatting may be missing. Original: https://www.theverge.com/policy/990267/microsoft-openai-new-york-times-authors-lawsuit