NewsCopyrightAI RegulationGenerative AI

Five Major Publishers Sue Meta Over Llama Training Data

Elsevier, Hachette, and three other publishers accuse Meta of using millions of copyrighted works without permission to train AI models. The dispute over fair use in AI development escalates.

5 major publishers sue Meta

Five Major Publishers Sue Meta Over Llama Training Data

Five of the world's largest book publishers have sued Meta Platforms and CEO Mark Zuckerberg. They accuse the company of using millions of copyrighted books and academic articles without authorization to train its Llama AI models. The lawsuit was filed on May 5 at the US District Court for the Southern District of New York – one of many similar cases now pending against AI developers.

Key Facts

  • Plaintiffs: Elsevier, Cengage Learning, Hachette Book Group, Macmillan Publishers, McGraw Hill, plus author Scott Turow and his company S.C.R.I.B.E. Inc
  • Allegation: Meta obtained copyrighted content from pirate libraries and illegal sources (including torrent downloads) and used it to train Llama
  • Demands: Damages, injunctive relief, and destruction of unauthorized copies of protected works
  • Status: Proposed class action, not yet certified by court

The Core Issue: Fair Use or Theft?

The publishers argue that Meta not only acted without licenses but deliberately removed or disregarded copyright management information and proceeded without seeking authorization. Zuckerberg is named personally – the plaintiffs claim he knew about and authorized decisions regarding the acquisition and use of training data.

Meta defends itself with the US fair-use doctrine: the company argues that using copyrighted material for AI training can be legal under certain conditions. Meta pledged to contest the lawsuit vigorously, citing earlier court decisions that treated some forms of AI training as "transformative use."

Training Data Volumes Explode

The dispute unfolds against a backdrop of rapidly expanding data consumption. Meta reported that Llama 3.1 was trained on over 15 trillion tokens. Llama 4 (April 2025) used up to 40 trillion tokens. For comparison: Chinese developer DeepSeek trained DeepSeek-V2 with 8.1 trillion tokens, later DeepSeek-V3 with 14.8 trillion.

Model Tokens Provider
Llama 3.1 >15 trillion Meta
Llama 4 up to 40 trillion Meta
DeepSeek-V3 14.8 trillion DeepSeek

A Precedent with Global Reach

This case is one of many: authors, publishers, news organizations, and artists are suing AI companies worldwide. The central question is everywhere the same: Is copying works for model training fair use – especially when materials come from illegal sources?

The answer will not be decided in New York alone. It could influence European proceedings, where copyright law is structured differently and the EU AI Act imposes additional requirements.

What This Means for German Companies

German AI developers and publishers should watch this case closely. It demonstrates that the question of training data legality is not abstract – it will be decided in court and can result in substantial damages. Anyone training AI models in Germany or Europe should be able to document where data comes from and whether licenses exist. The defense of "that was transformative use" could prove expensive if courts rule otherwise.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.