[{"data":1,"prerenderedAt":28},["ShallowReactive",2],{"nr-en-meta-verlage-klage-llama-trainingsdaten":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":10,"trackerLabel":10,"headlineStat":22,"image":23,"ogImage":24,"imageAlt":5,"csv":10,"minutes":25,"words":26,"html":27},"meta-verlage-klage-llama-trainingsdaten","Five Major Publishers Sue Meta Over Llama Training Data","Elsevier, Hachette, and three other publishers accuse Meta of using millions of copyrighted works without permission to train AI models. The dispute over fair use in AI development escalates.","2026-07-30","04:19","2026-07-30T04:19:00+02:00","","July 30, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Copyright","AI Regulation","Generative AI","Litigation","Training Data","5 major publishers sue Meta","\u002Fnewsroom\u002Fimg\u002Fmeta-verlage-klage-llama-trainingsdaten.webp","\u002Fog-nr\u002Fmeta-verlage-klage-llama-trainingsdaten.en.png",2,495,"\u003Cp>\u003Cstrong>Five of the world&#39;s largest book publishers have sued Meta Platforms and CEO Mark Zuckerberg.\u003C\u002Fstrong> They accuse the company of using millions of copyrighted books and academic articles without authorization to train its Llama AI models. The lawsuit was filed on May 5 at the US District Court for the Southern District of New York – one of many similar cases now pending against AI developers.\u003C\u002Fp>\n\u003Ch2>Key Facts\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Plaintiffs\u003C\u002Fstrong>: Elsevier, Cengage Learning, Hachette Book Group, Macmillan Publishers, McGraw Hill, plus author \u003Cstrong>Scott Turow\u003C\u002Fstrong> and his company S.C.R.I.B.E. Inc\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Allegation\u003C\u002Fstrong>: Meta obtained copyrighted content from \u003Cstrong>pirate libraries and illegal sources\u003C\u002Fstrong> (including torrent downloads) and used it to train Llama\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Demands\u003C\u002Fstrong>: Damages, injunctive relief, and \u003Cstrong>destruction of unauthorized copies\u003C\u002Fstrong> of protected works\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Status\u003C\u002Fstrong>: Proposed class action, not yet certified by court\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The Core Issue: Fair Use or Theft?\u003C\u002Fh2>\n\u003Cp>The publishers argue that Meta not only acted without licenses but deliberately \u003Cstrong>removed or disregarded copyright management information\u003C\u002Fstrong> and proceeded without seeking authorization. Zuckerberg is named personally – the plaintiffs claim he knew about and authorized decisions regarding the acquisition and use of training data.\u003C\u002Fp>\n\u003Cp>Meta defends itself with the \u003Cstrong>US fair-use doctrine\u003C\u002Fstrong>: the company argues that using copyrighted material for AI training can be legal under certain conditions. Meta pledged to contest the lawsuit vigorously, citing earlier court decisions that treated some forms of AI training as &quot;transformative use.&quot;\u003C\u002Fp>\n\u003Ch2>Training Data Volumes Explode\u003C\u002Fh2>\n\u003Cp>The dispute unfolds against a backdrop of rapidly expanding data consumption. Meta reported that Llama 3.1 was trained on \u003Cstrong>over 15 trillion tokens\u003C\u002Fstrong>. Llama 4 (April 2025) used up to \u003Cstrong>40 trillion tokens\u003C\u002Fstrong>. For comparison: Chinese developer DeepSeek trained DeepSeek-V2 with 8.1 trillion tokens, later DeepSeek-V3 with 14.8 trillion.\u003C\u002Fp>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Model\u003C\u002Fth>\n\u003Cth>Tokens\u003C\u002Fth>\n\u003Cth>Provider\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Llama 3.1\u003C\u002Ftd>\n\u003Ctd>&gt;15 trillion\u003C\u002Ftd>\n\u003Ctd>Meta\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Llama 4\u003C\u002Ftd>\n\u003Ctd>up to 40 trillion\u003C\u002Ftd>\n\u003Ctd>Meta\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>DeepSeek-V3\u003C\u002Ftd>\n\u003Ctd>14.8 trillion\u003C\u002Ftd>\n\u003Ctd>DeepSeek\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Ch2>A Precedent with Global Reach\u003C\u002Fh2>\n\u003Cp>This case is one of many: authors, publishers, news organizations, and artists are suing AI companies worldwide. The central question is everywhere the same: \u003Cstrong>Is copying works for model training fair use – especially when materials come from illegal sources?\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>The answer will not be decided in New York alone. It could influence European proceedings, where copyright law is structured differently and the EU AI Act imposes additional requirements.\u003C\u002Fp>\n\u003Ch2>What This Means for German Companies\u003C\u002Fh2>\n\u003Cp>German AI developers and publishers should watch this case closely. It demonstrates that the question of training data legality is not abstract – it will be decided in court and can result in substantial damages. Anyone training AI models in Germany or Europe should be able to document where data comes from and whether licenses exist. The defense of &quot;that was transformative use&quot; could prove expensive if courts rule otherwise.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fwww.nationthailand.com\u002Fbusiness\u002Fcorporate\u002F40069202\">Nation Thailand\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1785410832856]