Microsoft Submits Data to Claim AI Training on Copyrighted Material is Fair Use in Legal Battle with Publishers.
8
What is the Viqus Verdict?
We evaluate each news story based on its real impact versus its media hype to offer a clear and objective perspective.
AI Analysis:
Moderate current hype surrounding technical details, but the underlying legal challenge represents a genuinely high-impact decision point that could structurally change how all AI models are trained and deployed.
Article Summary
In a significant legal development, Microsoft has released analysis of 8.2 million Copilot chat logs to support its defense against copyright claims filed by major publishers like The New York Times and authors' guilds. The core of Microsoft's argument rests on quantitative data, which they interpret to show that the AI's output rarely regurgitates substantial copyrighted text, often containing fewer than 16 words from source material. Microsoft maintains that even when overlap occurs, the function of the resulting LLM is fundamentally transformative, constituting 'fair use.' This filing is a crucial component of a consolidated legal battle that questions the economic model of generative AI products built on ingested copyrighted works, suggesting that the current practices are legally sound despite the visible commercial conflicts.Key Points
- Microsoft presented data showing that even across millions of chat logs, the LLM rarely reproduces substantial or extended copyrighted passages, arguing the output is highly original.
- The company asserts that the act of using copyrighted material for training, and the resulting AI's functionality, constitutes a 'transformative' purpose that qualifies as fair use.
- This filing is a high-stakes attempt to secure a summary judgment, aiming to preemptively settle the legal challenges raised by major rights holders, including NYT and authors' guilds.

