Anthropic Sued Over Claude’s Song Lyrics Training Data

💡A major copyright case could reshape how AI developers source, filter, and document training data.
⚡ 30-Second TL;DR
What Changed
Sony Music Publishing and Warner Chappell filed a copyright lawsuit against Anthropic in California.
Why It Matters
The case could increase legal and financial risks for AI companies that train models on unlicensed copyrighted text. It may also push developers to strengthen dataset provenance, filtering, licensing, and memorization testing.
What To Do Next
Audit your training and fine-tuning corpora using Hugging Face dataset cards and provenance records, and remove or license copyrighted lyric datasets before deployment.
Key Points
- •Sony Music Publishing and Warner Chappell filed a copyright lawsuit against Anthropic in California.
- •The publishers allege Claude was trained on song lyrics obtained from pirate archives.
- •Dario Amodei and Benjamin Mann were named personally, with damages of up to $150,000 per composition sought.
- •A Munich court ruled in November 2025 that memorizing lyrics inside a model constitutes reproduction.
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •The lawsuit involves a coalition of 35 music-publishing entities, significantly expanding the scope beyond just Sony and Warner Chappell.
- •Plaintiffs are seeking an additional $25,000 per instance for the removal or alteration of copyright management information (CMI) in addition to the $150,000 per composition.
- •The legal complaint alleges that Anthropic utilized 'destructive scanning' of physical books as part of its data acquisition pipeline.
- •The current litigation leverages evidence unsealed during the 2025 'Bartz v. Anthropic' case, which concluded in a $1.5 billion settlement regarding pirated book data.
- •Anthropic is accused of scraping licensed lyric platforms like MusixMatch and LyricFind, rather than relying solely on pirate archives.
📊 Competitor Analysis▸ Show
| Feature | Anthropic (Claude) | OpenAI (GPT-4o) | Google (Gemini) |
|---|---|---|---|
| Copyright Litigation Status | High (Music/Books) | High (NYT/Authors) | Moderate (Class Actions) |
| Training Data Transparency | Low | Low | Low |
| Licensing Strategy | Defensive/Litigious | Emerging Partnerships | Emerging Partnerships |
🛠️ Technical Deep Dive
- The core technical allegation centers on 'verbatim output' where the model reproduces copyrighted lyrics, serving as evidence of memorization during the training phase.
- The training pipeline allegedly incorporated data from illicit repositories including Library Genesis and Pirate Library Mirror.
- The lawsuit claims the model architecture facilitates the retrieval of training data through specific prompting, challenging the notion that models only perform probabilistic inference.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.



