๐Ÿค–Stalecollected in 27m

How to access the Books3 dataset for research

PostLinkedIn
๐Ÿค–Read original on Reddit r/MachineLearning

๐Ÿ’กUnderstand the current accessibility of the controversial Books3 dataset used in major LLM training.

โšก 30-Second TL;DR

What Changed

Books3 is a large-scale collection of books used for training LLMs.

Why It Matters

The availability of Books3 significantly impacts how researchers train models on long-form text. Legal restrictions on this dataset may force a shift toward licensed or public domain data sources.

What To Do Next

Consult your organization's legal counsel regarding the use of scraped datasets like Books3 to avoid potential copyright liability in your training pipeline.

Who should care:Researchers & Academics

Key Points

  • โ€ขBooks3 is a large-scale collection of books used for training LLMs.
  • โ€ขThe dataset is currently subject to legal scrutiny regarding copyright infringement.
  • โ€ขResearchers are seeking legitimate access paths amidst widespread takedowns.

๐Ÿง  Deep Insight

AI-generated analysis for this event โ€” not the original article.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขBooks3 is a subset of The Pile, a 800GB dataset curated by EleutherAI, which was specifically designed for training large language models.
  • โ€ขThe dataset was removed from public repositories like The Eye and Hugging Face in 2023 following a DMCA takedown notice issued by the Authors Guild.
  • โ€ขLegal proceedings involving Books3 include high-profile class-action lawsuits, such as Silverman et al. v. OpenAI, which allege that the dataset contains copyrighted works used without authorization.
  • โ€ขResearchers often face significant ethical and legal hurdles when attempting to access Books3, as many mirrors are now hosted on non-indexed or decentralized platforms to evade copyright enforcement.
  • โ€ขThe controversy surrounding Books3 has catalyzed a shift in the AI industry toward 'data transparency' initiatives, with some organizations now prioritizing the use of public domain or licensed datasets for model training.

๐Ÿ› ๏ธ Technical Deep Dive

  • Books3 consists of approximately 196,640 books in plain text format.
  • The dataset was originally scraped from Bibliotik, a private torrent tracker, and compiled into a single corpus.
  • It is primarily utilized for pre-training transformer-based architectures to improve long-range dependency modeling and narrative coherence.
  • The data is typically processed into tokenized formats (e.g., BPE or SentencePiece) before being fed into LLM training pipelines.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Standardized licensing will replace 'scraping-first' data collection.
Ongoing litigation is forcing AI companies to adopt rigorous data provenance standards to mitigate legal liability.
Publicly available 'Books3-like' datasets will become increasingly rare.
Increased legal scrutiny and the risk of copyright infringement claims are driving developers to favor proprietary or licensed data sources.

โณ Timeline

2020-12
EleutherAI releases The Pile, which includes the Books3 dataset.
2023-08
The Authors Guild and various authors file lawsuits citing the use of Books3 in AI training.
2023-09
Books3 is removed from The Eye and Hugging Face following DMCA takedown requests.
2024-02
A federal judge dismisses some copyright claims against Meta regarding Books3 but allows others to proceed.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.