AI Labs Have Used Aggressive Data Collection Tactics

Tech firms have faced scrutiny for scraping copyrighted content and bypassing paywalls to train large language models.

Updated on Oct. 7, 2026 in Artificial Intelligence

Bold flat-color editorial illustration featuring a thick bundle of fiber optic cables, representing institutional data harvesting in an abstract style.
AI laboratories face increasing backlash over data collection practices, including scraping copyrighted content and bypassing paywalls to fuel large language models. AI Illustration. Upload story photo >

Live Poll

Should AI companies be required to obtain explicit consent before training on your public content?

Artificial intelligence laboratories have employed large-scale data collection strategies that include bypassing paywalls and scraping millions of copyrighted documents. These practices have prompted content creators to implement widespread restrictions on automated web access.

Why it matters

The open web no longer provides enough high-quality data to improve model capabilities, forcing labs to aggressively source content while facing rising backlash over labor and copyright concerns.

AI companies utilize massive datasets like Common Crawl, which contained over 2 million documents from nytimes.com and 91,000 works from multiple outlets. Currently, 45 percent of C4 data faces restrictions due to terms-of-service blocks.

The players

Brent Hecht

He is the Director of Applied Science at Microsoft who has publicly criticized the data collection practices of AI laboratories.

Dataset Providers Alliance

This coalition released a position paper in 2026 advocating for an opt-in system regarding the use of web data for model training.

Google

This technology company reportedly pays Reddit roughly $60 million annually to license user data for its platforms.

OpenAI

This AI research organization spends approximately $70 million annually to secure access to proprietary data for training its models.

The details

Internal projects like Project Mango and Project Taxi funneled news content for training, while researchers noted that some news organizations saw click-through rates decline by 51 to 94 percent following AI product launches. Microsoft Director of Applied Science Brent Hecht has criticized these data practices, describing them as an astonishing theft of labor.

Timeline

  1. Shutterstock reported $138 million in data licensing revenue for 2024.

  2. The Dataset Providers Alliance formed in the summer of 2026.

  3. An MIT-led study on data provenance was published in October 2026.

  4. The AI training dataset market is projected to multiply by the early 2030s.

The Tech Race

This development follows the pattern set by the Dataset Providers Alliance position paper, which challenges the industry's reliance on open-web scraping. The shift toward licensing deals represents a move away from unregulated data ingestion toward structured, paid partnerships.

Users may encounter more paywalls or restricted access as publishers shield content from AI scrapers. Additionally, content creators might see changes to how their work is credited or compensated as the market for licensed training data matures.

The takeaway

The tension between AI developers and content creators signals a permanent change in how web information is valued and protected. Organizations should expect a transition toward exclusive data-licensing agreements as the legal and ethical standards for model training solidify.

Further reading

For more information on the evolving landscape of data rights, visit our Artificial Intelligence section.

Source note: This article includes information reported by WebProNews.

Live Poll

Should AI companies be required to obtain explicit consent before training on your public content?