Biphoo News

collapse
Home / Daily News Analysis / Sony Music and Warner Chappell sue Anthropic over song lyrics in Claude's training data

Sony Music and Warner Chappell sue Anthropic over song lyrics in Claude's training data

Aug 31, 2026  Twila Rosenbaum  5 views
Sony Music and Warner Chappell sue Anthropic over song lyrics in Claude's training data

Sony Music Publishing and Warner Chappell have filed a lawsuit against Anthropic in a Northern California court, alleging that the AI company engaged in a “brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale” to train its Claude models. The publishers have named Anthropic CEO Dario Amodei and Benjamin Mann, a co-founder, personally in the complaint, according to Business Insider.

The lawsuit centers on song lyrics that the publishers claim were taken from pirate archives, specifically Library Genesis and Pirate Library Mirror. These are the same shadow library repositories that were at the heart of a $1.5 billion settlement Anthropic reached with authors earlier. The complaint lists well-known works, including Eye of the Tiger, Hallelujah, September, Livin’ On a Prayer, and Great Balls of Fire, along with compositions by Mariah Carey and Taylor Swift.

The publishers are demanding a jury trial and statutory damages, seeking up to $150,000 for each composition used in training. That figure is the statutory ceiling for willful infringement under U.S. copyright law, not an amount any court has already awarded. It is an opening demand in a case that has yet to be answered.

The contrast between this demand and the previous settlement is striking. In the earlier authors’ case, each title earned about $3,000, split between the author and their publisher, leaving roughly $1,500 per side. The gap between those numbers — $150,000 per work versus $3,000 per title — illustrates the entire negotiating landscape. One figure represents what two sides agreed upon in a settlement; the other is a maximalist opening position in a lawsuit that could take years to resolve.

Previous Legal Troubles and the Shadow Library Connection

Anthropic has faced copyright litigation before. The company previously agreed to a $1.5 billion settlement with a group of authors who accused it of using pirated books from Library Genesis and similar sources to train its models. Those archives, which host millions of copyrighted works without authorization, have become a recurring point of tension between AI developers and rightsholders.

The new lawsuit extends that battle into the music industry. The plaintiffs argue that Anthropic’s training data included vast amounts of song lyrics scraped from these pirate libraries, and that the resulting AI models are capable of reproducing those lyrics verbatim when prompted. The complaint alleges that this practice was not accidental but systematic, with the company knowingly using pirated sources to build a competitive commercial product.

Legal experts note that the personal naming of Amodei and Mann is an aggressive tactic. Under U.S. copyright law, corporate officers can sometimes be held personally liable for infringement if they had the ability to supervise the infringing activity and a financial interest in it. The plaintiffs are likely attempting to pressure the executives individually, though such claims are often difficult to prove.

European Courts Have Already Weighed In

While this case is unfolding in the United States, a European court has already addressed a similar question involving song lyrics and AI training. The Regional Court of Munich ruled against OpenAI in November 2025, finding that memorizing lyrics inside a model constitutes reproduction under German and EU copyright law. The court also held that outputs reciting those lyrics amount to communication to the public, a separate exclusive right of rightsholders.

Importantly, the Munich court found that the text and data mining (TDM) exception did not apply. The exception allows some forms of automated analysis of copyrighted works, but the court concluded that permanent memorization goes beyond transient analysis. In this case, the rightsholder had also opted out of TDM, which triggered the exception’s opt-out provision. The judgment is not final and could be appealed, but it signals a shift in how European courts view AI training practices.

The EU’s TDM exception carries a second condition that is highly relevant to the Anthropic lawsuit. It applies only to works to which the miner had lawful access. A pirate library is never lawful access. This means that even if the TDM exception were otherwise applicable, using content scraped from Library Genesis or Pirate Library Mirror would likely void that defense. The same logic could influence U.S. courts, though American copyright law does not have a comparable general TDM exception.

The AI Act and Transparency Obligations

On top of the court ruling, the European Union’s AI Act adds another layer of regulatory pressure. General-purpose model providers, including companies like Anthropic and OpenAI, must maintain a copyright policy and publish a summary of the training data they use. This requirement is enforced by a dedicated unit in Brussels, which can investigate complaints and issue corrective actions.

This creates a striking asymmetry between the two sides of the Atlantic. American rightsholders often have to go to court and engage in discovery to find out what training data was used and whether their works were included. In contrast, European rightsholders are entitled to be told, at least in summary form, what data went into a model. The AI Act’s transparency provisions are not yet fully operational, but they represent a fundamental shift toward proactive disclosure.

For music publishers, the ability to know exactly which lyrics were used and how is crucial. The U.S. litigation process offers a path to that information through discovery, but it is expensive, slow, and often obscured by claims of trade secrets. The AI Act’s summary requirement could provide a faster route to identifying infringement, at least for models distributed in the EU.

Implications for the AI and Music Industries

The outcome of this case could reshape how AI companies approach training data, particularly when it comes to creative content like song lyrics. If the publishers succeed in proving willful infringement, the statutory damages could be enormous — potentially billions of dollars if hundreds of thousands of compositions were involved. But even a mid-range award would send a strong signal.

The music industry has been increasingly vigilant about AI’s use of copyrighted material. Major labels and publishers have formed alliances with AI companies that license content properly, such as the deals between Universal Music Group and various AI startups. However, cases like this one demonstrate that unauthorized use remains a significant concern. The fact that Anthropic allegedly used pirate libraries rather than licensed sources is particularly damaging in the court of public opinion, regardless of the legal outcome.

For Anthropic, the lawsuit arrives at a time when the company is already under scrutiny for its data practices. The $1.5 billion settlement with authors was one of the largest ever in an AI copyright case, and this new suit suggests that the underlying issue has not been fully resolved. The personal naming of the CEO and a co-founder could also complicate fundraising and business development, as investors may view legal exposure as a risk factor.

The Munich ruling adds a cross-border dimension. Even if Anthropic were to win in the U.S., it still faces potential liability in Europe, where the legal landscape is less favorable to AI companies. The AI Act’s enforcement mechanisms mean that compliance with copyright laws is not just a matter of litigation but also of regulatory oversight. Companies that fail to publish adequate summaries or maintain copyright policies could face fines of up to 3% of global annual turnover.

There is also the question of how AI models actually memorize lyrics. Researchers have shown that large language models can reproduce long sequences from training data, especially when the data is duplicated many times. Pirate libraries often contain the same song lyrics in multiple files, which increases the likelihood of memorization. This technical reality undercuts the argument that AI models do not copy but merely learn patterns.

The case is still in its early stages. Anthropic has not yet filed a response. The publishers are asking for a jury trial, which could be years away. In the meantime, the AI industry will be watching closely, as the interpretation of copyright law in this context could set precedents for future lawsuits involving other types of creative works, from news articles to visual art.

The asymmetry between U.S. and EU approaches is unlikely to resolve quickly. American courts rely on case-by-case adjudication, while European regulators are building a proactive enforcement framework. For rightsholders, the European path may offer more immediate answers. For AI companies, the patchwork of legal obligations creates uncertainty and risk. This lawsuit is just one front in a wider war over the intersection of artificial intelligence and copyright law — a conflict that will likely define the next decade of digital creativity.


Source: TNW | Anthropic News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy