Times Filed Unredacted Microsoft and OpenAI Communications
New court documents allege that AI companies bypassed news paywalls to train generative models.
Updated on Oct. 3, 2026 in Artificial Intelligence

Live Poll
Should AI companies be required to compensate news publishers for using their content to train models?
The New York Times has submitted unredacted internal communications from Microsoft and OpenAI executives as evidence in their ongoing legal battle. The filings allege that AI companies utilized unauthorized methods to bypass news paywalls for training data.
Why it matters
The disclosures highlight how AI companies rely on copyrighted news content to train models, raising questions about the future of digital journalism. Internal documents suggest executives were aware of these concerns, with some labeling the disruption an existential threat to publishers.
OpenAI employees allegedly employed technical hacks to bypass publisher paywalls while training large language models. These models demonstrate an ability to predict and reproduce text from copyrighted articles.
The players
The New York Times
This major American news organization is a prominent publisher currently engaged in litigation against leading artificial intelligence firms.
OpenAI
This artificial intelligence research organization is the developer of the ChatGPT platform and large language models central to the current copyright dispute.
Microsoft
A multinational technology corporation that provides the cloud infrastructure and significant financial investment for OpenAI products.
Brent Hecht
He is an executive whose internal communications regarding AI training practices were made public through the recent court filing.
Greg Brockman
A co-founder of OpenAI who acknowledged the technical capability of models to predict text from specific news source materials.
The details
Internal documents reveal that Microsoft and OpenAI staff identified the harvesting of news content as a potential risk to the economic sustainability of publishers. Execs acknowledged an AI content strategy described as a doom loop that threatens the employment of human data generators.
Timeline
October 2, 2026: Unredacted court filings were publicly reported.
The Tech Race
This dispute marks a critical juncture in the ongoing copyright litigation between news publishers and generative AI developers. It highlights the tension between rapid innovation in large language models and the intellectual property rights of the content creators upon which they train.
The outcome of this case may redefine how users access free news content online if publishers implement stricter technical barriers to prevent AI scraping. Readers might eventually see changes in digital subscription access or shifts in how information is summarized by AI tools.
The takeaway
These filings reveal that the tension between generative AI development and journalism involves documented internal concerns about content sustainability. Stakeholders should anticipate heightened legal scrutiny regarding how tech companies source training data from the open web.
Further reading
For broader context on the industry's shift toward regulation, see the latest updates in Artificial Intelligence.
More information
Review the full details in the unredacted court filing documents.
Source note: This article includes information reported by The Corvallis Advocate.
Live Poll
Should AI companies be required to compensate news publishers for using their content to train models?










