Five of the world's largest book publishers have sued Meta Platforms and CEO Mark Zuckerberg. They accuse the company of using millions of copyrighted books and academic articles without authorization to train its Llama AI models. The lawsuit was filed on May 5 at the US District Court for the Southern District of New York – one of many similar cases now pending against AI developers.
Key Facts
- Plaintiffs: Elsevier, Cengage Learning, Hachette Book Group, Macmillan Publishers, McGraw Hill, plus author Scott Turow and his company S.C.R.I.B.E. Inc
- Allegation: Meta obtained copyrighted content from pirate libraries and illegal sources (including torrent downloads) and used it to train Llama
- Demands: Damages, injunctive relief, and destruction of unauthorized copies of protected works
- Status: Proposed class action, not yet certified by court
The Core Issue: Fair Use or Theft?
The publishers argue that Meta not only acted without licenses but deliberately removed or disregarded copyright management information and proceeded without seeking authorization. Zuckerberg is named personally – the plaintiffs claim he knew about and authorized decisions regarding the acquisition and use of training data.
Meta defends itself with the US fair-use doctrine: the company argues that using copyrighted material for AI training can be legal under certain conditions. Meta pledged to contest the lawsuit vigorously, citing earlier court decisions that treated some forms of AI training as "transformative use."
Training Data Volumes Explode
The dispute unfolds against a backdrop of rapidly expanding data consumption. Meta reported that Llama 3.1 was trained on over 15 trillion tokens. Llama 4 (April 2025) used up to 40 trillion tokens. For comparison: Chinese developer DeepSeek trained DeepSeek-V2 with 8.1 trillion tokens, later DeepSeek-V3 with 14.8 trillion.
| Model | Tokens | Provider |
|---|---|---|
| Llama 3.1 | >15 trillion | Meta |
| Llama 4 | up to 40 trillion | Meta |
| DeepSeek-V3 | 14.8 trillion | DeepSeek |
A Precedent with Global Reach
This case is one of many: authors, publishers, news organizations, and artists are suing AI companies worldwide. The central question is everywhere the same: Is copying works for model training fair use – especially when materials come from illegal sources?
The answer will not be decided in New York alone. It could influence European proceedings, where copyright law is structured differently and the EU AI Act imposes additional requirements.
What This Means for German Companies
German AI developers and publishers should watch this case closely. It demonstrates that the question of training data legality is not abstract – it will be decided in court and can result in substantial damages. Anyone training AI models in Germany or Europe should be able to document where data comes from and whether licenses exist. The defense of "that was transformative use" could prove expensive if courts rule otherwise.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




