Sony and Warner Chappell Sue Anthropic Over Alleged Pirated Training Data for Claude
Sony Music Publishing and Warner Chappell Music have sued Anthropic and two co-founders, alleging copyrighted works were obtained and copied in connection with Claude training.
Summary
Sony Music Publishing, Warner Chappell Music and affiliated publishers have sued Anthropic and co-founders Dario Amodei and Benjamin Mann, alleging large-scale copyright infringement connected to data used to develop Claude.
The complaint was filed on August 28, 2026, in the U.S. District Court for the Northern District of California as case 5:26-cv-09217. It alleges that copyrighted works were obtained through torrenting, scraping and other acquisition methods, then copied during model training and, in some instances, reproduced through Claude outputs.
These are allegations in a newly filed lawsuit, not findings by the court. Anthropic told TechCrunch that it disagrees with the publishers’ claims and intends to defend itself in court.
Quick Take
- Sony Music Publishing and Warner Chappell are now direct plaintiffs against Anthropic, widening the lab’s music-copyright exposure.
- The complaint also names Dario Amodei and Benjamin Mann and seeks damages, injunctive relief and a jury trial.
- Requested statutory ceilings reach $150,000 per willfully infringed work and $25,000 per alleged CMI violation; no damages have been awarded.
- The case targets not only AI training and outputs but also how source material was allegedly acquired.
- The practical issue for AI companies is increasingly dataset provenance: where training data came from, under what rights, and with what records.
What the publishers allege
The complaint describes several distinct forms of alleged conduct rather than a single copyright theory. The publishers say Anthropic and its founders participated in or benefited from large-scale acquisition of copyrighted material through torrenting, scraping and downloading, including material allegedly obtained from pirate-library sources and lyric services.
They further allege that protected works were copied during the process of training Claude and that Claude can produce verbatim or near-verbatim lyrics from some copyrighted songs. A separate claim concerns copyright management information (CMI) — identifying information associated with protected works — which the publishers say was removed or altered during data processing.
The filing asserts four principal causes of action:
| Claim | Defendants targeted | What is being alleged |
|---|---|---|
| Direct infringement through torrenting | Anthropic, Amodei and Mann | Unauthorized acquisition and copying of protected works |
| Contributory infringement through torrenting | Amodei and Mann | Personal involvement in or contribution to alleged infringing acquisition |
| Direct copyright infringement | Anthropic | Copying protected works in connection with Claude development and use |
| Removal or alteration of CMI | Anthropic | Removal or alteration of copyright-management information |
None of those claims has yet been proven in this lawsuit.
The key legal distinction: acquisition is not the same question as training
Much of the generative-AI copyright debate has focused on whether using copyrighted material to train a model can qualify as fair use. This complaint attacks an earlier step as well: how the copies used for training were allegedly obtained in the first place.
That distinction matters because the legal analysis does not necessarily collapse into one question. An AI developer could have one argument about whether model training is transformative while still facing a separate dispute over whether the source copies themselves were lawfully acquired.
Why it matters: The risk is moving upstream. AI companies increasingly need to defend not only what a model outputs, but also the provenance of the corpus behind it.
For developers and AI labs, this turns dataset governance from a back-office documentation task into a potential legal and commercial control.
Five risk layers the case puts under scrutiny
The Sony-Warner complaint is useful as a map of how AI copyright exposure can accumulate across a development pipeline.
| Risk layer | What this case puts under scrutiny |
|---|---|
| Data acquisition | Alleged torrenting, scraping and unauthorized downloading |
| Model training | Alleged copying of protected works as training inputs |
| Model outputs | Alleged verbatim or near-verbatim lyric reproduction |
| Copyright metadata | Alleged removal or alteration of CMI |
| Individual liability | Claims naming company founders personally |
These layers matter because a favorable ruling for an AI company on one issue would not automatically resolve all the others.
What the damages numbers mean — and what they do not
The publishers ask for statutory damages of up to $150,000 for each work found to have been willfully infringed and up to $25,000 for each qualifying CMI violation. They also seek injunctive relief and other remedies.
Those numbers can produce very large theoretical exposure when multiplied across many works, but they are requested statutory ceilings, not an amount Anthropic has been ordered to pay.
Before any damages award, the court would have to resolve core questions including which works are actionable, whether infringement occurred, whether any infringement was willful, which defendants are liable and what remedies are legally available.
That is why headline descriptions of a “multi-billion-dollar” case should be read as estimates of potential exposure under the plaintiffs’ theory — not as a current judgment against Anthropic.
Why dataset provenance is becoming a business issue
The broader implication goes beyond Anthropic and the music industry. Frontier AI developers increasingly need a defensible record of where training material originated, how it was accessed, what rights applied and whether attribution or copyright metadata was preserved.
That can matter in litigation, but also in enterprise procurement, licensing negotiations, partnerships, due diligence and future regulation. A company that can reconstruct the lineage of its training corpus has a different risk profile from one that cannot.
Music is especially important here because licensing structures are mature. If litigation keeps pushing liability toward acquisition and provenance, pressure may increase for licensed training-data deals rather than large-scale collection first and legal reconciliation later.
What to watch next
The immediate signal will be Anthropic’s formal court response. After that, the important procedural questions include whether Anthropic moves to dismiss parts of the case, how the court treats the claims against Amodei and Mann personally, and whether acquisition claims are analyzed separately from training-copy and output claims.
It will also be worth watching how this lawsuit interacts with Anthropic’s other copyright disputes and whether music publishers increasingly pursue licensing negotiations in parallel with litigation.
AI World Scope Take
The most important part of this case is not the largest damages number. It is the increasingly explicit separation between how copyrighted material enters an AI corpus, what happens to it during training, and what the resulting model can reproduce.
For years, the public AI-copyright debate was often reduced to one question: Is training fair use? This litigation shows why that framing may be too narrow.
The strategic question for AI companies is becoming: Can you prove that the data behind your models was assembled through defensible acquisition, licensing and governance practices?
If courts continue to treat acquisition, training, outputs and metadata as distinct layers, there may never be one ruling that settles the entire AI-training copyright debate at once. Dataset provenance may instead become a durable part of how AI companies manage legal risk.
Sources & Documentation
Sources used for this article, with source type and publisher shown where available.
- documentationComplaint — Sony Music Publishing (US) LLC et al. v. Anthropic PBC et al.Visit Source
- documentationSony Music Publishing (US) LLC et al. v. Anthropic PBC et al. — DocketVisit Source
- newsSony Music, Warner sue Anthropic, alleging a “brazen campaign” of intellectual property theftVisit Source
- newsSony Music Publishing and Warner Chappell sue Anthropic in multi-billion dollar lawsuitVisit Source