[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-laion-80-millionen-videos-ki-forschung":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"laion-80-millionen-videos-ki-forschung","LAION Releases 80 Million Videos for Open AI Research","With the Big Video Dataset, LAION launches one of the largest open video collections – ten million hours of material for AI labs worldwide.","2026-08-29","13:31","2026-08-29T13:31:00+02:00","","August 29, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Open-Source AI","Video Datasets","AI Research","LAION","Copyright","\u002Fstand-der-ki","Progress at the AI Frontier","80 million videos, ten million hours","\u002Fnewsroom\u002Fimg\u002Flaion-80-millionen-videos-ki-forschung.webp","\u002Fog-nr\u002Flaion-80-millionen-videos-ki-forschung.en.png",2,408,"\u003Cp>LAION has released the \u003Cstrong>Big Video Dataset (BVD)\u003C\u002Fstrong>, creating one of the largest freely accessible video databases for AI research. The scale is impressive: \u003Cstrong>80 million videos\u003C\u002Fstrong> totaling \u003Cstrong>ten million hours\u003C\u002Fstrong> of footage, from which \u003Cstrong>55 million clips\u003C\u002Fstrong> with automatically generated descriptions have been extracted. Add \u003Cstrong>300 million individual frames\u003C\u002Fstrong> to that. The dataset is made available to researchers worldwide and enables them to train video-language models.\u003C\u002Fp>\n\u003Ch2>Quick Facts\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>80 million videos\u003C\u002Fstrong> sourced from 1.3 billion URLs (primarily YouTube, English-language)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Performance gain\u003C\u002Fstrong>: Models trained on BVD outperform the previous reference dataset InternVid by up to \u003Cstrong>2.1 percentage points\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Legal basis\u003C\u002Fstrong>: 2024 Hamburg court ruling permits LAION to collect copyrighted material for non-commercial research\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Access\u003C\u002Fstrong>: Free for research purposes; commercial use excluded\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>How the Training Works\u003C\u002Fh2>\n\u003Cp>The BVD&#39;s strength lies in its multimodal approach: models learn simultaneously from \u003Cstrong>video, audio, and text\u003C\u002Fstrong>. The system understands not just which visual content matches which descriptions, but also which sounds and music accompany them. This combination drives the measured performance improvements on standard video-text benchmarks.\u003C\u002Fp>\n\u003Ch2>Legal Gray Area\u003C\u002Fh2>\n\u003Cp>LAION operates in sensitive territory here. The organization relies on a \u003Cstrong>2024 Hamburg Regional Court ruling\u003C\u002Fstrong> that permits collecting copyrighted material for non-commercial research. This is an important precedent for open-source AI in Germany and Europe – though also contested. LAION explicitly asks users to &quot;respect the rights and copyrights of content creators.&quot; Whether this legal position will hold up in higher courts or when commercial applications emerge remains uncertain.\u003C\u002Fp>\n\u003Ch2>What This Means for European Research\u003C\u002Fh2>\n\u003Cp>For German and European AI labs, BVD is a strategic asset. Until now, large video datasets for training video-language models were hard to access – researchers not working at OpenAI, Google, or other US giants relied on smaller or licensed sources. With BVD, universities, research institutes, and European startups can now work at the frontier of video AI models without depending on proprietary APIs. This matters especially for applications in medicine, Industry 4.0, or accessibility, where European data protection standards and local requirements count.\u003C\u002Fp>\n\u003Cp>However: the dataset is primarily English-language and YouTube-centric. Those wanting to train on German or European content will need additional sources. The question of how long the legal foundation holds if commercial applications emerge also remains relevant for long-term research planning.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fbig-video-dataset-laion-stellt-zehn-millionen-stunden-videomaterial-fuer-offene-ki-forschung-bereit\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Flaion-drops-massive-open-video-dataset-with-10-million-hours-of-footage-for-ai-research\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1788003441039]