[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-xiaomi-robotics-1-daten-schlagen-modellgroesse":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"xiaomi-robotics-1-daten-schlagen-modellgroesse","Xiaomi-Robotics-1: Why More Data Beats Bigger Models","Xiaomi trained an AI model for robots that proves data volume matters more than compute power. With over 100,000 hours of movement data, success rates jumped from 25 to 75 percent—and gains keep climbing.","2026-07-21","11:25","2026-07-21T11:25:00+02:00","","July 21, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Robotics AI","Machine Learning","Data Training","Xiaomi","Automation","\u002Fstand-der-ki","AI Progress","Success rate increased from 25 % to 75 %","\u002Fnewsroom\u002Fimg\u002Fxiaomi-robotics-1-daten-schlagen-modellgroesse.webp","\u002Fog-nr\u002Fxiaomi-robotics-1-daten-schlagen-modellgroesse.en.png",3,549,"\u003Cp>Xiaomi has achieved a breakthrough in robotics AI that challenges conventional wisdom about model training. The company trained its \u003Cstrong>Xiaomi-Robotics-1\u003C\u002Fstrong> model on over \u003Cstrong>100,000 hours of movement data\u003C\u002Fstrong> and discovered: more training data improves performance far more dramatically than larger models with greater compute power. Success rates on unfamiliar tasks climbed from roughly 25 to 75 percent—and researchers report this trend hasn&#39;t plateaued yet.\u003C\u002Fp>\n\u003Ch2>Key Facts\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>100,000+ hours\u003C\u002Fstrong> of movement data collected from over \u003Cstrong>1,700 different environments\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>Data volume improves performance \u003Cstrong>significantly more\u003C\u002Fstrong> than model size\u003C\u002Fli>\n\u003Cli>Success rate on unfamiliar tasks: \u003Cstrong>25 % to 75 %\u003C\u002Fstrong> increase\u003C\u002Fli>\n\u003Cli>Handheld grippers with cameras used instead of expensive robots for data collection\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The Data Problem in Robotics\u003C\u002Fh2>\n\u003Cp>Robotics AI faces a fundamental challenge that language models don&#39;t: while LLMs can train on vast portions of the public internet, useful training data for robot movements is extremely scarce. Typically, humans must manually guide each robot through every movement—a slow, expensive process that produces repetitive data from identical tasks in identical settings.\u003C\u002Fp>\n\u003Cp>Xiaomi solved this creatively: instead of deploying actual robots, the team used portable \u003Cstrong>handheld grippers with attached cameras\u003C\u002Fstrong> that people simply pick up and operate by hand. This allowed researchers to record manipulation tasks in kitchens, offices, stores, factory floors, and outdoor spaces without a robot present at all. The result: over 100,000 hours of movement recordings.\u003C\u002Fp>\n\u003Ch2>Automated Labeling as a Scaling Solution\u003C\u002Fh2>\n\u003Cp>A dataset this large created a new problem: each recording needed a text description for the model to learn from. Manual labeling was impractical. Xiaomi deployed another AI model to automatically describe each movement segment. The team labeled the entire dataset in approximately \u003Cstrong>two weeks\u003C\u002Fstrong>.\u003C\u002Fp>\n\u003Cp>Xiaomi then transferred training to actual robots—wheeled models and dual-arm systems. To account for differences between handheld grippers and robot arms, the team combined its own recordings from real apartments with open-source robotics datasets and annotated handheld data.\u003C\u002Fp>\n\u003Ch2>Data Beats Compute Power\u003C\u002Fh2>\n\u003Cp>Tests were conclusive: a larger model improved performance, but \u003Cstrong>more training data delivered dramatically larger gains than more compute power\u003C\u002Fstrong>. Researchers concluded that progress in robotics AI depends primarily on collecting larger and more diverse datasets.\u003C\u002Fp>\n\u003Cp>This pattern differs from the classical rule for language models, where model size and data volume should grow at roughly equal rates. With visual data—and apparently with robot movements—additional data helps far more than additional compute.\u003C\u002Fp>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Metric\u003C\u002Fth>\n\u003Cth>Result\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Success rate (before)\u003C\u002Ftd>\n\u003Ctd>~25 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Success rate (after)\u003C\u002Ftd>\n\u003Ctd>~75 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Training data\u003C\u002Ftd>\n\u003Ctd>100,000+ hours\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Environments\u003C\u002Ftd>\n\u003Ctd>1,700+\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Labeling time\u003C\u002Ftd>\n\u003Ctd>~2 weeks\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Ch2>What This Means for Companies\u003C\u002Fh2>\n\u003Cp>This finding carries weight for robotics and automation firms globally. The conventional assumption that you need more expensive, larger AI models is being challenged here—instead, the focus should shift to \u003Cstrong>data collection and diversity\u003C\u002Fstrong>. For companies training robots, this could mean investments in data gathering (such as handheld systems) might be more efficient than purchasing additional compute power. At the same time, questions remain about how these insights transfer to specialized, highly complex robotics tasks—and whether a 75 percent success rate meets production requirements.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fxiaomi-robotics-1-shows-that-more-data-beats-bigger-models-when-training-robots-to-move\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1784638331758]