NewsBreakingType: news

Sony and Warner Chappell Sue Anthropic Over Alleged Pirated Training Data for Claude

Sony Music Publishing and Warner Chappell Music have sued Anthropic and two co-founders, alleging copyrighted works were obtained and copied in connection with Claude training.

AW
AI World Scope Editorial DeskSource-backed editorial coverage
August 30, 20266 min read
AI World Scope
AI World Scope diagram showing the lawsuit’s three contested layers: alleged data acquisition, Claude training, and outputs plus copyright metadata.

Summary

Sony Music Publishing, Warner Chappell Music and affiliated publishers have sued Anthropic and co-founders Dario Amodei and Benjamin Mann, alleging large-scale copyright infringement connected to data used to develop Claude.

The complaint was filed on August 28, 2026, in the U.S. District Court for the Northern District of California as case 5:26-cv-09217. It alleges that copyrighted works were obtained through torrenting, scraping and other acquisition methods, then copied during model training and, in some instances, reproduced through Claude outputs.

These are allegations in a newly filed lawsuit, not findings by the court. Anthropic told TechCrunch that it disagrees with the publishers’ claims and intends to defend itself in court.

Quick Take

  • Sony Music Publishing and Warner Chappell are now direct plaintiffs against Anthropic, widening the lab’s music-copyright exposure.
  • The complaint also names Dario Amodei and Benjamin Mann and seeks damages, injunctive relief and a jury trial.
  • Requested statutory ceilings reach $150,000 per willfully infringed work and $25,000 per alleged CMI violation; no damages have been awarded.
  • The case targets not only AI training and outputs but also how source material was allegedly acquired.
  • The practical issue for AI companies is increasingly dataset provenance: where training data came from, under what rights, and with what records.

What the publishers allege

The complaint describes several distinct forms of alleged conduct rather than a single copyright theory. The publishers say Anthropic and its founders participated in or benefited from large-scale acquisition of copyrighted material through torrenting, scraping and downloading, including material allegedly obtained from pirate-library sources and lyric services.

They further allege that protected works were copied during the process of training Claude and that Claude can produce verbatim or near-verbatim lyrics from some copyrighted songs. A separate claim concerns copyright management information (CMI) — identifying information associated with protected works — which the publishers say was removed or altered during data processing.

The filing asserts four principal causes of action:

ClaimDefendants targetedWhat is being alleged
Direct infringement through torrentingAnthropic, Amodei and MannUnauthorized acquisition and copying of protected works
Contributory infringement through torrentingAmodei and MannPersonal involvement in or contribution to alleged infringing acquisition
Direct copyright infringementAnthropicCopying protected works in connection with Claude development and use
Removal or alteration of CMIAnthropicRemoval or alteration of copyright-management information

None of those claims has yet been proven in this lawsuit.

The key legal distinction: acquisition is not the same question as training

Much of the generative-AI copyright debate has focused on whether using copyrighted material to train a model can qualify as fair use. This complaint attacks an earlier step as well: how the copies used for training were allegedly obtained in the first place.

That distinction matters because the legal analysis does not necessarily collapse into one question. An AI developer could have one argument about whether model training is transformative while still facing a separate dispute over whether the source copies themselves were lawfully acquired.

Why it matters: The risk is moving upstream. AI companies increasingly need to defend not only what a model outputs, but also the provenance of the corpus behind it.

For developers and AI labs, this turns dataset governance from a back-office documentation task into a potential legal and commercial control.

Five risk layers the case puts under scrutiny

The Sony-Warner complaint is useful as a map of how AI copyright exposure can accumulate across a development pipeline.

Risk layerWhat this case puts under scrutiny
Data acquisitionAlleged torrenting, scraping and unauthorized downloading
Model trainingAlleged copying of protected works as training inputs
Model outputsAlleged verbatim or near-verbatim lyric reproduction
Copyright metadataAlleged removal or alteration of CMI
Individual liabilityClaims naming company founders personally

These layers matter because a favorable ruling for an AI company on one issue would not automatically resolve all the others.

What the damages numbers mean — and what they do not

The publishers ask for statutory damages of up to $150,000 for each work found to have been willfully infringed and up to $25,000 for each qualifying CMI violation. They also seek injunctive relief and other remedies.

Those numbers can produce very large theoretical exposure when multiplied across many works, but they are requested statutory ceilings, not an amount Anthropic has been ordered to pay.

Before any damages award, the court would have to resolve core questions including which works are actionable, whether infringement occurred, whether any infringement was willful, which defendants are liable and what remedies are legally available.

That is why headline descriptions of a “multi-billion-dollar” case should be read as estimates of potential exposure under the plaintiffs’ theory — not as a current judgment against Anthropic.

Why dataset provenance is becoming a business issue

The broader implication goes beyond Anthropic and the music industry. Frontier AI developers increasingly need a defensible record of where training material originated, how it was accessed, what rights applied and whether attribution or copyright metadata was preserved.

That can matter in litigation, but also in enterprise procurement, licensing negotiations, partnerships, due diligence and future regulation. A company that can reconstruct the lineage of its training corpus has a different risk profile from one that cannot.

Music is especially important here because licensing structures are mature. If litigation keeps pushing liability toward acquisition and provenance, pressure may increase for licensed training-data deals rather than large-scale collection first and legal reconciliation later.

What to watch next

The immediate signal will be Anthropic’s formal court response. After that, the important procedural questions include whether Anthropic moves to dismiss parts of the case, how the court treats the claims against Amodei and Mann personally, and whether acquisition claims are analyzed separately from training-copy and output claims.

It will also be worth watching how this lawsuit interacts with Anthropic’s other copyright disputes and whether music publishers increasingly pursue licensing negotiations in parallel with litigation.

AI World Scope Take

The most important part of this case is not the largest damages number. It is the increasingly explicit separation between how copyrighted material enters an AI corpus, what happens to it during training, and what the resulting model can reproduce.

For years, the public AI-copyright debate was often reduced to one question: Is training fair use? This litigation shows why that framing may be too narrow.

The strategic question for AI companies is becoming: Can you prove that the data behind your models was assembled through defensible acquisition, licensing and governance practices?

If courts continue to treat acquisition, training, outputs and metadata as distinct layers, there may never be one ruling that settles the entire AI-training copyright debate at once. Dataset provenance may instead become a durable part of how AI companies manage legal risk.

Sources & Documentation

Sources used for this article, with source type and publisher shown where available.

  • documentationComplaint — Sony Music Publishing (US) LLC et al. v. Anthropic PBC et al.
    Visit Source
  • documentationSony Music Publishing (US) LLC et al. v. Anthropic PBC et al. — Docket
    Visit Source
  • newsSony Music, Warner sue Anthropic, alleging a “brazen campaign” of intellectual property theft
    Visit Source
  • newsSony Music Publishing and Warner Chappell sue Anthropic in multi-billion dollar lawsuit
    Visit Source
AI World Scope Briefing

Stay ahead in AI

Join the list for selected AI news, model releases, comparisons and tool updates when new briefings are published.

Your email is stored for AI World Scope briefing delivery.