ViqusViqus
Navigate
Company
Blog
About Us
Contact
System Status
Enter Viqus Hub

Microsoft Submits Data to Claim AI Training on Copyrighted Material is Fair Use in Legal Battle with Publishers.

Microsoft Copilot copyright law AI training fair use New York Times legal battle LLM
September 04, 2026
Source: The Verge AI
Viqus Verdict Logo Viqus Verdict Logo 8
The Foundational Lawsuit for Generative AI
Media Hype 6/10
Real Impact 8/10

Article Summary

In a significant legal development, Microsoft has released analysis of 8.2 million Copilot chat logs to support its defense against copyright claims filed by major publishers like The New York Times and authors' guilds. The core of Microsoft's argument rests on quantitative data, which they interpret to show that the AI's output rarely regurgitates substantial copyrighted text, often containing fewer than 16 words from source material. Microsoft maintains that even when overlap occurs, the function of the resulting LLM is fundamentally transformative, constituting 'fair use.' This filing is a crucial component of a consolidated legal battle that questions the economic model of generative AI products built on ingested copyrighted works, suggesting that the current practices are legally sound despite the visible commercial conflicts.

Key Points

  • Microsoft presented data showing that even across millions of chat logs, the LLM rarely reproduces substantial or extended copyrighted passages, arguing the output is highly original.
  • The company asserts that the act of using copyrighted material for training, and the resulting AI's functionality, constitutes a 'transformative' purpose that qualifies as fair use.
  • This filing is a high-stakes attempt to secure a summary judgment, aiming to preemptively settle the legal challenges raised by major rights holders, including NYT and authors' guilds.

Why It Matters

This is not just a technical dispute; it is a foundational legal battle defining the economic viability and legal limits of the entire generative AI industry. If a judge rules in favor of Microsoft’s 'transformative use' argument, it provides a massive, potentially industry-wide legal green light for using copyrighted data for training foundational models, significantly de-risking the entire AI sector for large tech players. Conversely, a ruling against them could force a painful overhaul of how companies ingest data, potentially necessitating expensive licensing agreements and stifling rapid foundational model development. Professionals must track this outcome as it sets the precedent for digital content usage globally.

You might also be interested in