AI Labs Faced Scrutiny Over Data Scraping Tactics
Researchers and publishers are clashing over the aggressive acquisition of web content used to train new AI models.
Updated on Oct. 7, 2026 in Artificial Intelligence

Live Poll
Should AI companies be required to obtain explicit consent before training on your public content?
AI laboratories have come under fire for allegedly using aggressive strategies to scrape high-quality training data, including bypassing paywalls and site restrictions. As high-quality content becomes scarce, major companies are increasingly signing multimillion-dollar licensing deals to secure the data needed for model growth.
Why it matters
The race to develop more capable large language models has led to a data drought, prompting companies to source information from protected or proprietary channels. This shift has ignited a debate over the ethics of data usage and the financial impact on news organizations facing declining traffic.
AI companies have faced criticism as 25 percent of high-quality web sources now block crawlers, while news outlets report traffic declines between 51 and 94 percent following AI product launches. Meanwhile, firms like OpenAI and Google pay Reddit $70 million and $60 million annually for data access.
The players
Brent Hecht
He is a Microsoft Director of Applied Science who has publicly characterized current AI training practices as theft.
Dataset Providers Alliance
This organization released a position paper advocating for an opt-in system for the use of data in AI training.
The technology giant pays Reddit approximately $60 million annually to access data for its AI development.
OpenAI
The research laboratory reportedly pays Reddit about $70 million per year for continued data access.
Shutterstock
The stock media provider generated $138 million in revenue during 2024 through the licensing of its data for AI.
The details
AI developers have reportedly utilized internal initiatives like Project Mango and Project Taxi to funnel content across organizations for training purposes. This practice has led to datasets containing over 2 million documents from specific news sites, fueling claims from industry experts that these methods amount to data theft.
Timeline
Shutterstock reported $138 million in data licensing revenue for 2024.
The Dataset Providers Alliance formed in the summer of 2026.
The term slurp juice was coined in October 2026.
The Tech Race
The current reliance on aggressive data scraping stands in opposition to the opt-in framework recently proposed in the Dataset Providers Alliance position paper. This struggle signals a transition from the era of free, open-web scraping to a closed-loop market where data is treated as a premium asset.
As publishers tighten restrictions, users may encounter more aggressive paywalls and login requirements on news sites attempting to protect their content from crawlers. Additionally, readers might see fewer free, high-quality search results as publishers prioritize internal licensing over public access.
The takeaway
The tension between AI labs and content providers suggests that the era of 'free' training data is rapidly ending. Consumers should expect web content to become increasingly siloed behind private licensing deals as firms seek to secure stable data supplies.
Further reading
For more background on the industry, visit Artificial Intelligence.
Source note: This article includes information reported by WebProNews.
Live Poll
Should AI companies be required to obtain explicit consent before training on your public content?










