Sony and Warner have sued Anthropic over song lyrics, sheet music and derivative works, in a 48-page complaint that reaches past the company to name Amodei personally: it alleges he directed, approved and oversaw the infringement carried out by Mann and other Anthropic employees. The publishers call it "one of the largest and most blatant ongoing thefts of intellectual property in history." The suit arrives roughly a year after Anthropic paid $1.5 billion to settle a near-identical claim from book authors, and it aims at the same weak point that settlement exposed: not what Anthropic trained on, but how it obtained the files.
The publishers want damages for each work infringed, and, separately, damages for the removal of copyright management information — the attribution notices and other identifying data attached to a work. The statutory ceilings are $150,000 per infringed work and $25,000 per removal.
That September 2025 settlement was the largest copyright settlement in US history, and its shape matters for reading this one. Anthropic did not pay because it used protected material to train a model. It paid because of how the material was acquired: downloaded illegally over torrents.
Sony and Warner point at the same conduct. Anthropic, the complaint says, torrented at least seven million books from the pirate libraries LibGen and PiLiMi. The plaintiffs treat each of those downloads as a standalone act of infringement — whether or not the work ever reached a commercial Claude model.
That framing is the cleverest thing in the filing. It severs the question of liability from the hardest question in AI copyright law, which is what a model does with a text once it has ingested it. If the download is the wrong, there is no need to argue about weights, memorisation or fair use at all. Anthropic has already conceded, with money, that this theory works.
The rest of the complaint widens the target. Anthropic is accused of scraping lyrics from the licensed platforms MusixMatch and LyricFind without publisher consent, in breach of those services' terms of use. The plaintiffs also challenge Anthropic's use of the Books3, The Pile and Common Crawl datasets, which they say contain unlicensed material, and accuse the company of scanning and destroying second-hand songbooks and sheet music collections. Lumping destructive scanning of physical copies in with torrenting is a choice: it signals that the publishers are contesting the whole acquisition pipeline, not only the unambiguously illegal part of it.
The most consequential section concerns what happened after the downloads. Anthropic has publicly denied training commercial Claude models directly on books from LibGen and PiLiMi. The plaintiffs argue that denial turns entirely on what Anthropic means by "training." At least one commercial Claude model, they say, was trained on synthetic data generated by a non-commercial model — and that non-commercial model was trained on LibGen and PiLiMi texts. They further allege Anthropic used such a model to produce reinforcement feedback for a commercial Claude. The full scope of the claim is left to discovery.
This is the part worth watching, and it is not really about music. Every frontier lab now trains on synthetic data produced by its own internal models, and "we did not train on that corpus" is the standard industry answer to every acquisition question. If a court accepts that a pirated corpus travels through a teacher model into a student model, that answer stops working — not just for Anthropic, and not just for lyrics. The plaintiffs have essentially proposed that distillation can launder provenance, and asked a judge to say it cannot.
Notably absent from everything described in the complaint is the number that determines what Anthropic is actually exposed to. Seven million is a count of books. This is a music case, and the $150,000 ceiling applies per musical work infringed. How many compositions the publishers are asserting is not stated, and until it is, the headline arithmetic everyone will reach for — seven million times $150,000 — is the wrong multiplication entirely.
There is also a competing theory of the case already on the books elsewhere. In November 2025 the Munich Regional Court held that copyrighted lyrics remain reproductions even inside a model's parameters, and that a model emitting them constitutes unlawful public disclosure. The court put responsibility for the output on the model's operator rather than on the user who wrote the prompt — even where the prompt was deliberately engineered to extract protected lyrics.
The two theories pull in opposite directions, and only one of them is survivable. Acquisition-based liability makes the download the crime, which is expensive but finite: Anthropic has demonstrated it can write that cheque. The Munich reading makes the weights themselves the infringing copy, and there is no version of that where a shipped model is clean.