SourceStalecollected in 36m

AI models rely on human-curated Wikipedia data

Read original on 虎嗅
#data-curation#training-data#human-in-the-loop

Discover the hidden human labor and data curation processes that power the world's most advanced AI models.

30-Second TL;DR

What Changed

AI models use Wikipedia as a primary 'hippocampus' for structured, reliable knowledge.

Why It Matters

The sustainability of high-quality training data is at risk if the human-in-the-loop ecosystem is disrupted by AI-driven traffic shifts.

What To Do Next

When building datasets, prioritize human-in-the-loop verification processes to ensure the quality and reliability of your model's knowledge base.

Who should care:Researchers & Academics

Key Points

  • •AI models use Wikipedia as a primary 'hippocampus' for structured, reliable knowledge.
  • •1% of 'super-editors' contribute the vast majority of Wikipedia's content, forming the bedrock of AI training data.
  • •AI companies are increasingly using commercial APIs (Wikimedia Enterprise) to access this data.
  • •Generative AI is reducing human traffic to Wikipedia, potentially threatening the sustainability of the volunteer ecosystem.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.