Introduction
Recent unsealed court documents have reignited debate over the ethics and legality of how artificial intelligence companies train their models on publicly available and paywalled content. The New York Times' ongoing lawsuit against OpenAI and Microsoft has brought to light internal communications that label AI data scraping as theft, quantify its impact on newsrooms, and detail alleged efforts to circumvent paywalls.
What Happened
According to newly revealed filings, a Microsoft internal memo from January 2023, authored by Applied Science director Brent Hecht, described AI data scraping as the largest theft of labor in human history. The same documents show that OpenAI leadership acknowledged that chatbot outputs could substitute for visiting original news sites, potentially undermining the very publishers that trained the models. The unsealed material also details how companies allegedly bypassed paywalls, stripped copyright notices from training sets, and assembled datasets containing tens of thousands of New York Times articles alongside millions more from other outlets.
- Internal Microsoft records show Copilot's answer engine reduced click-through rates for The New York Times by up to 93% compared to traditional search.
- OpenAI leadership warned that publishers face significant threats from chatbots that largely substitute for original content.
- Microsoft CEO Satya Nadella testified that any paywalled content used for AI training should be licensed, and that he would demand model retraining if unauthorized scraping were discovered.
- The filings reveal that OpenAI's mid-training datasets alone contained over 91,692 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting, plus more than 2 million documents from nytimes.com.
- Allegedly, OpenAI employees discussed hacks to bypass The New York Times paywall, and the company built training datasets by mass-scraping Common Crawl and other open web repositories.
Why This Matters
The case strikes at the core of the AI industry's current training paradigm. If courts determine that large-scale scraping infringes copyright or harms market value, it could force a fundamental shift in how AI models are built, potentially increasing costs and limiting capabilities. For publishers, the stakes are equally high: a ruling against them could further erode revenue streams, while a ruling in their favor might establish new licensing frameworks and compensation models. The outcome will likely shape AI development, digital publishing, and copyright law for years to come.
Key Takeaways
- AI companies' training practices are under unprecedented scrutiny following unsealed court filings.
- Microsoft and OpenAI executives privately characterized large-scale data scraping as theft and acknowledged its disruptive potential.
- Internal data suggests AI-powered answer engines can dramatically reduce traffic to original news sources.
- The lawsuit exposes a gap between AI industry norms and traditional copyright expectations around paywalled content.
- Scale of copying revealed: tens of thousands of NYT articles and millions of web pages appear in mid-training datasets.
Conclusion
As the legal battle between The New York Times, OpenAI, and Microsoft progresses, the tech and publishing worlds are watching closely. The unsealed filings offer a rare glimpse into the internal attitudes and practices that drive today's generative AI systems. Whether the courts will reshape the boundaries of fair use, or whether new legislation will emerge to govern AI training data, one thing is clear: the relationship between AI development and content creators will be a defining issue of the decade.




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.