AI Firms Buy Thousands of Books for Training

💡Book-buying and mass scanning expose the copyright and provenance risks behind frontier-model training data.
⚡ 30-Second TL;DR
What Changed
A Chinese AI company ordered 3,000 used books from a Dutch bookstore.
Why It Matters
AI developers may face greater legal and reputational risks when training models on digitized books or other scraped content. The revelations could accelerate demand for licensed corpora, provenance tracking, and clearer disclosure of training-data sources.
What To Do Next
Record every corpus source in a Hugging Face Dataset Card, including license, acquisition method, and deletion or retention policy, before using it for training.
Key Points
- •A Chinese AI company ordered 3,000 used books from a Dutch bookstore.
- •Anthropic reportedly scanned and discarded several million books.
- •The activity raises questions about copyright, licensing, and transparency in AI training-data procurement.
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •The Dutch bookstore involved, 'De Slegte,' reported that the bulk order was specifically requested to be in good condition, suggesting a preference for high-quality physical text for OCR (Optical Character Recognition) processing.
- •Anthropic's practice of scanning physical books is part of a broader industry trend to bypass digital copyright restrictions by utilizing 'analog' copies that may fall under different legal interpretations of fair use.
- •The procurement of physical books for AI training is often outsourced to third-party data brokers who specialize in acquiring and digitizing large-scale physical archives to avoid direct corporate liability.
- •Legal experts note that the 'discarding' of books after scanning creates a potential loophole in copyright law, as the AI firm does not retain the physical copy, complicating claims of unauthorized reproduction.
- •This practice highlights a shift in AI training strategies toward 'high-quality' data, as models increasingly suffer from 'model collapse' when trained on synthetic or low-quality internet-scraped data.
🛠️ Technical Deep Dive
- Data ingestion pipelines for physical media utilize high-speed industrial scanners equipped with automated page-turning technology to maximize throughput.
- OCR (Optical Character Recognition) engines used in these pipelines often employ specialized transformer-based models to correct errors in text extraction from aged or damaged physical pages.
- Pre-processing involves cleaning noise, deskewing images, and applying layout analysis to distinguish between body text, headers, and metadata before tokenization.
- The resulting datasets are often stored in proprietary formats optimized for rapid sequential access during the pre-training phase of Large Language Models (LLMs).
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本) ↗