Japanese Books Shipped Abroad for AI Training
Used bookstores report large-volume orders from overseas entities to harvest content for artificial intelligence.
Updated on Sept. 29, 2026 in Artificial Intelligence

Live Poll
Should companies be allowed to use copyrighted books to train artificial intelligence models?
Used bookstores across Japan have received mysterious, high-volume orders for non-fiction titles. Booksellers suspect these books are being destroyed during scanning to train artificial intelligence models.
Why it matters
The systematic procurement and destruction of printed works highlight the growing demand for vast, high-quality data sets to fuel the development of large language models. This trend raises concerns regarding the longevity of physical archives and the ethics of data collection for proprietary technology.
A major Japanese book distributor exported over 50 tons of books to the United States last year. Individual Tokyo stores have reported single transactions exceeding 1,000 volumes, with many shipments routed through a logistics center in Okayama Prefecture.
The players
Anthropic
This is an American artificial intelligence research company that develops the Claude chatbot and confirmed it purchases books for AI training purposes.
The details
Booksellers report that orders often arrive at regular intervals for non-fiction genres, including medicine, philosophy, and history. After the books arrive at designated facilities, they are typically scanned for data ingestion and destroyed, a process Anthropic has confirmed is used to train its Claude chatbot.
Timeline
A distributor sold 50 tons of books to the U.S. between 2025 and 2026.
Anthropic revealed its book-buying practices earlier in 2026.
The Tech Race
The reliance of AI models on physical books as training data is exemplified by the procurement processes used for the Claude AI chatbot. This trend marks a shift where physical media is being repurposed as raw data fodder to maintain competitiveness in the escalating development of large language models.
While this trend primarily affects the availability of rare or specific non-fiction titles in used bookstores, it signals a broader shift in how public knowledge is harvested. Readers may see changes in the availability of physical academic or reference texts as large-scale data ingestion continues.
The takeaway
The transformation of printed non-fiction into digital training sets illustrates the intense pressure companies face to secure proprietary data. Collectors and readers should be aware that physical books are increasingly being viewed as expendable assets in the global race for AI supremacy.
Further reading
Learn more about the infrastructure behind these systems in our Artificial Intelligence section.
Source note: This article includes information reported by Japan Today.
Live Poll
Should companies be allowed to use copyrighted books to train artificial intelligence models?







