NewsRobotics AIMachine LearningData Training

Xiaomi-Robotics-1: Why More Data Beats Bigger Models

Xiaomi trained an AI model for robots that proves data volume matters more than compute power. With over 100,000 hours of movement data, success rates jumped from 25 to 75 percent—and gains keep climbing.

Success rate increased from 25 % to 75 %

Xiaomi-Robotics-1: Why More Data Beats Bigger Models

Xiaomi has achieved a breakthrough in robotics AI that challenges conventional wisdom about model training. The company trained its Xiaomi-Robotics-1 model on over 100,000 hours of movement data and discovered: more training data improves performance far more dramatically than larger models with greater compute power. Success rates on unfamiliar tasks climbed from roughly 25 to 75 percent—and researchers report this trend hasn't plateaued yet.

Key Facts

  • 100,000+ hours of movement data collected from over 1,700 different environments
  • Data volume improves performance significantly more than model size
  • Success rate on unfamiliar tasks: 25 % to 75 % increase
  • Handheld grippers with cameras used instead of expensive robots for data collection

The Data Problem in Robotics

Robotics AI faces a fundamental challenge that language models don't: while LLMs can train on vast portions of the public internet, useful training data for robot movements is extremely scarce. Typically, humans must manually guide each robot through every movement—a slow, expensive process that produces repetitive data from identical tasks in identical settings.

Xiaomi solved this creatively: instead of deploying actual robots, the team used portable handheld grippers with attached cameras that people simply pick up and operate by hand. This allowed researchers to record manipulation tasks in kitchens, offices, stores, factory floors, and outdoor spaces without a robot present at all. The result: over 100,000 hours of movement recordings.

Automated Labeling as a Scaling Solution

A dataset this large created a new problem: each recording needed a text description for the model to learn from. Manual labeling was impractical. Xiaomi deployed another AI model to automatically describe each movement segment. The team labeled the entire dataset in approximately two weeks.

Xiaomi then transferred training to actual robots—wheeled models and dual-arm systems. To account for differences between handheld grippers and robot arms, the team combined its own recordings from real apartments with open-source robotics datasets and annotated handheld data.

Data Beats Compute Power

Tests were conclusive: a larger model improved performance, but more training data delivered dramatically larger gains than more compute power. Researchers concluded that progress in robotics AI depends primarily on collecting larger and more diverse datasets.

This pattern differs from the classical rule for language models, where model size and data volume should grow at roughly equal rates. With visual data—and apparently with robot movements—additional data helps far more than additional compute.

Metric Result
Success rate (before) ~25 %
Success rate (after) ~75 %
Training data 100,000+ hours
Environments 1,700+
Labeling time ~2 weeks

What This Means for Companies

This finding carries weight for robotics and automation firms globally. The conventional assumption that you need more expensive, larger AI models is being challenged here—instead, the focus should shift to data collection and diversity. For companies training robots, this could mean investments in data gathering (such as handheld systems) might be more efficient than purchasing additional compute power. At the same time, questions remain about how these insights transfer to specialized, highly complex robotics tasks—and whether a 75 percent success rate meets production requirements.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.