🗾Freshcollected in 53m

AI Firms Buy Thousands of Books for Training

AI Firms Buy Thousands of Books for Training
PostLinkedIn
🗾Read original on ITmedia AI+ (日本)

💡Book-buying and mass scanning expose the copyright and provenance risks behind frontier-model training data.

⚡ 30-Second TL;DR

What Changed

A Chinese AI company ordered 3,000 used books from a Dutch bookstore.

Why It Matters

AI developers may face greater legal and reputational risks when training models on digitized books or other scraped content. The revelations could accelerate demand for licensed corpora, provenance tracking, and clearer disclosure of training-data sources.

What To Do Next

Record every corpus source in a Hugging Face Dataset Card, including license, acquisition method, and deletion or retention policy, before using it for training.

Who should care:Researchers & Academics

Key Points

  • A Chinese AI company ordered 3,000 used books from a Dutch bookstore.
  • Anthropic reportedly scanned and discarded several million books.
  • The activity raises questions about copyright, licensing, and transparency in AI training-data procurement.

🧠 Deep Insight

AI-generated analysis for this event.

🔑 Enhanced Key Takeaways

  • The Dutch bookstore involved, 'De Slegte,' reported that the bulk order was specifically requested to be in good condition, suggesting a preference for high-quality physical text for OCR (Optical Character Recognition) processing.
  • Anthropic's practice of scanning physical books is part of a broader industry trend to bypass digital copyright restrictions by utilizing 'analog' copies that may fall under different legal interpretations of fair use.
  • The procurement of physical books for AI training is often outsourced to third-party data brokers who specialize in acquiring and digitizing large-scale physical archives to avoid direct corporate liability.
  • Legal experts note that the 'discarding' of books after scanning creates a potential loophole in copyright law, as the AI firm does not retain the physical copy, complicating claims of unauthorized reproduction.
  • This practice highlights a shift in AI training strategies toward 'high-quality' data, as models increasingly suffer from 'model collapse' when trained on synthetic or low-quality internet-scraped data.

🛠️ Technical Deep Dive

  • Data ingestion pipelines for physical media utilize high-speed industrial scanners equipped with automated page-turning technology to maximize throughput.
  • OCR (Optical Character Recognition) engines used in these pipelines often employ specialized transformer-based models to correct errors in text extraction from aged or damaged physical pages.
  • Pre-processing involves cleaning noise, deskewing images, and applying layout analysis to distinguish between body text, headers, and metadata before tokenization.
  • The resulting datasets are often stored in proprietary formats optimized for rapid sequential access during the pre-training phase of Large Language Models (LLMs).

🔮 Future ImplicationsAI analysis grounded in cited sources

Increased regulation of physical book digitization for AI training.
Legislators are likely to introduce 'right to digitize' amendments that specifically address the mass scanning of copyrighted physical works by AI entities.
Rise of 'Data Provenance' certification for AI models.
To avoid legal risks, enterprise AI customers will increasingly demand transparency reports detailing the exact origin and licensing status of the training data used in models.

Timeline

2023-07
Anthropic faces initial copyright lawsuits regarding the use of copyrighted books in training data.
2024-03
Reports emerge detailing the use of 'Books3' and other large-scale book datasets in major LLM training.
2025-02
Anthropic expands data acquisition strategies to include diverse physical and digital archives to improve reasoning capabilities.
2026-05
Dutch bookstore 'De Slegte' fulfills the bulk order for 3,000 books, sparking public discourse on AI data procurement.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本)