Skip to content

Sony and Warner Sue Anthropic Over Training Data Piracy

4 min read

Sony and Warner Sue Anthropic Over Training Data Piracy
Photo by KATRIN BOLOVTSOVA on Pexels

This Lawsuit Has a Precedent — and It Cost Anthropic $1.5 Billion

Sony Music Publishing and Warner Chappell Music filed a 48-page complaint against Anthropic in California federal court on August 29, 2026, alleging the company ran a “brazen campaign” of torrenting copyrighted song lyrics and sheet music to train Claude. The suit names Dario Amodei and Benjamin Mann personally and seeks up to $150,000 per infringed copyright.

What distinguishes this case is the legal ground it stands on. In June 2025, a federal court ruled on a related Anthropic case brought by book authors: AI training itself was fair use, but downloading over seven million books from piracy sites to build that training corpus was not. That distinction — lawful training, unlawful acquisition — is exactly what Sony and Warner invoke here. Anthropic settled that case for $1.5 billion.

What the Complaint Alleges

The complaint alleges two acquisition channels. First, Anthropic downloaded lyrics via torrent from Library Genesis and the Pirate Library Mirror — the same repositories cited in the 2025 book case. Second, the company scraped licensed platforms including Musixmatch and LyricFind, both of which pay licensing fees to publishers for the very content allegedly scraped.

The named works span five decades: “Ain’t No Mountain High Enough,” “All I Want for Christmas Is You,” “Eye of the Tiger,” “Hallelujah,” Taylor Swift’s “Cruel Summer” and “Paper Rings,” and hundreds more. The complaint covers “tens of thousands” of compositions — deliberate ambiguity that keeps the damages ceiling open. Publishers also allege Claude can reproduce lyrics verbatim on request, which they say they tested personally.

The Legal Landscape After the 2025 Ruling

The 2025 ruling created a bifurcated defense for Anthropic. On training, it has favorable precedent: the four-factor fair use test weighed in its favor because training is transformative, models don’t reproduce works wholesale, and no market substitution was proven. On acquisition, the precedent runs against it: downloading pirated content is infringement regardless of downstream use. Obtaining a pirated song is infringement even if you only listen once. The downstream use does not launder the upstream conduct.

Anthropic’s internal data-acquisition records — surfaced during discovery — will determine whether the publishers can prove their case. The specific piracy sites named are identical to those in the 2025 book case, which suggests the publishers have seen documents pointing there.

Why Music Is a Harder Case Than Books

Licensed lyric databases — Musixmatch, LyricFind, Genius — generate real revenue for publishers. If Claude returns song lyrics verbatim, users bypass those services rather than visiting them. That is a direct substitution claim — the factor courts weigh most heavily in fair use analysis. A person checking whether a Taylor Swift lyric is “shake it off” or “shake it out” may accept Claude’s answer instead of clicking through to a licensed site. This argument is structurally stronger than the book-publisher claims, where readers do not accept model reconstructions in place of actual novels.

What a Settlement Would Normalize

Anthropic’s response is four words: “We intend to defend.” The company raised capital at a $65 billion valuation earlier this year, so it can litigate seriously. But the 2025 book settlement shows it will negotiate when exposure is large enough. “Tens of thousands” of works at $150,000 each produces a theoretical ceiling in the billions.

Universal Music Group — which sued Anthropic in 2023 and again in early 2026 — is watching what Sony and Warner extract before finalizing its own position. If this case settles with per-track licensing fees for training data, every other rights holder in AI litigation will demand the same terms. For developers building or fine-tuning models, the message is now unambiguous: training data provenance is a legal surface. The fair use defense covers transformative training. It does not cover torrenting.

Further Reading

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Enjoyed this? Get one AI insight per day.

Join engineers and decision-makers who start their morning with vortx.ch. No fluff, no hype — just what matters in AI.