Anthropic thought it had settled its book-piracy woes with a $1.5 billion payout. Music publishers, however, are not impressed, and on Friday, they filed a lawsuit that makes the AI company's 'historic settlement' look less like a penance and more like a cover charge for a party that never ended.

The music industry heavyweights - including Sony, EMI, and Warner Chappell - are alleging that Anthropic's 'brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale' didn't stop at books. Oh no, it extended to 'thousands upon thousands' of their copyrighted musical compositions, from Beatles songbooks to Taylor Swift's 'best' songs to 'VH1's 100 Greatest Songs of Rock & Roll.'

According to the lawsuit, Anthropic's 'mass campaign' began in July 2021 when co-founder Benjamin Mann personally used BitTorrent to download millions of pirated books from Library Genesis (LibGen). CEO Dario Amodei allegedly approved the torrenting, and both are named as defendants. When LibGen was shut down by the FBI, pirates simply copied it to create Z-Library. When that got shut down, Anthropic allegedly obtained a copy of the copy via the 'Pirate Library Mirror' (PiLiMi). Mann reportedly messaged colleagues that the mirror dropped 'just in time!' - to which a staffer replied, 'zlibrary my beloved.'

These internal messages, revealed in the book authors' case, are being used as evidence of Anthropic's 'extolling' of piracy. The publishers claim that Anthropic's torrenting included at least hundreds of books containing sheet music and song lyrics, and they intend to reveal the 'full extent' of it through discovery.

Anthropic denies using the pirated books to train its commercial Claude models, but publishers argue that 'training' can be defined broadly. They point to a 'pretraining' phase where Anthropic may have used synthetic data from non-commercial models that were trained on LibGen or PiLiMi text to reinforce commercial Claude models. And then there's the guardrail issue: unsealed documents allegedly show Anthropic continued using the LibGen dataset to check for output similarity, even after stopping LLM training on it.

Anthropic's spokesperson dismissed the suit as 'the third lawsuit from the same lawyers, recycling allegations from cases already before the courts,' and asserted that 'training generative AI models is a transformative fair use - as the court held in Bartz.' But publishers note that the fair use ruling in Bartz hinged on a lack of proven market harm - something they think they can demonstrate.

The publishers argue that AI-generated songs are already topping music charts, competing directly with the human songwriters whose unpaid work allegedly trained the very AI displacing them. They cite the US Copyright Office's observation that outputs competing in a market for a type of work can dilute royalty pools. They also allege that Claude was intentionally trained to regurgitate lyrics, even when not requested, and can reproduce the 'heart' of popular songs on demand.

Amodei's testimony from the book case is also dredged up: he said Anthropic could have legally purchased copyrighted works but torrented them instead to avoid a 'legal/practice/business slog.' The court summarized it as downloading pirated books 'to avoid the trouble of paying for them.' The publishers emphasize that Anthropic never approached them for licenses, unlike its AI rivals.

Anthropic closely guards its training data sources, but publishers want the court to force transparency: an accounting of training data, methods, and capabilities. Otherwise, they warn, songwriters will find it harder to make a living, and the market will be flooded with low-quality AI copies of recognizable songs. And, of course, publishers won't get paid for licensing songs to train AI. As the complaint puts it, Anthropic 'enriches itself through the uncompensated exploitation of Music Publishers' and their songwriters' labor.'