At the World Robot Conference (WRC 2026) in Beijing’s Yizhuang this August, humanoid robots sorted packages on conveyor belts and tightened screws at workstations—such “getting-things-done” scenarios have become as commonplace as theatrical demonstrations. But beneath this bright line of “deployment,” a hidden thread about data is quietly determining the next phase of competition in the embodied intelligence industry.
Unitree Robotics founder Wang Xingxing said something at the conference that nearly every vendor privately acknowledges: “However much data you have—especially high-quality data—that’s how much AI capability you have.” Humanoid robots’ limbs can already run, jump, and do backflips, but their “brains” are still starving. One frequently cited statistic: China currently has only 500,000 hours of compliant real-world physical interaction data, while commercial robot deployment requires tens of millions of hours—a gap exceeding 99%.
From hundreds of thousands to tens of millions, the difference spans two orders of magnitude. So players at WRC each showed their hand: some open-sourced human behavior data, some sold data-collection gloves, some bet on simulation, and some insisted on real-machine, real-world data. Beyond the noise, a few embodied intelligence companies with no booths and little attention are holding the industry’s most scarce data.
The Open-Source Camp: Turning Data Collection into a “Mass Movement”
Lightwheel Intelligence made one of the biggest data-related moves at this year’s WRC. On August 20, the company released EgoSuite-Open100K, the world’s first 100,000-hour-class full-modal open-source human behavior dataset, covering 7 major environment categories, 128 scenario types, over 15,000 collection scenes, and more than 15,000 tasks. The data is primarily captured from a head-mounted first-person perspective, with some wrist cameras added, providing hand pose, full-body pose, and event-level semantic annotations. The first batch is already available for download via Hugging Face.
Lightwheel Intelligence CEO Yang Haibo explained that relying solely on real-machine teleoperation can hardly support training data supply at the ten-million-hour level or above. The industry urgently needs two parallel paths: human video data and simulation-synthesized data. The company proposed a five-year, 10-billion-hour embodied intelligence data co-construction plan—in plain terms, “I can’t collect it all myself, so let the whole industry collect together.”
The Real-Data Camp: Data Costs Aren’t as High as You Think
Xinghai Tu CEO Gao Jiyang insists on prioritizing real data. He ran the numbers: after factoring in data costs, compute costs, and R&D engineer labor costs, “data costs are actually relatively manageable—not as high as people imagine. The most expensive part is actually R&D engineers’ time.” Xinghai Tu’s real-data assets include the GOD real-scene dataset open-sourced in September 2025, which ranks first globally in downloads; expanded data sources through investments in companies like Jianzhi and Yuanliu; and a plan to scale real data to 1 million hours by 2026. At the same time, Xinghai Tu hasn’t abandoned simulation—its paper citations rank first in the world-model field.
Even Unitree Robotics, while using real robot operational data to help robots further adapt to the physical world, cannot avoid pre-training on massive internet data. Walking on two legs—one for “broad exposure,” one for “grounded reality”—will remain the norm for embodied intelligence for a long time.
Extreme Operating Conditions: Data No Other Method Can Collect
All the data from the players above—whether human behavior video, simulation synthesis, or real-machine teleoperation—comes almost entirely from relatively standardized, structured, safe, and controllable scenarios. But there is one category of data that none of these methods can capture: real-world physical interaction data under extreme operating conditions—mines 1,000 meters underground, beside smelting furnaces at over a thousand degrees Celsius, at the edge of blast zones in open-pit mines.
These are scenarios humans don’t want to enter (63% of mine accidents are caused by human error), simulation cannot model (the physical complexity of unstructured extreme conditions exceeds the modeling capability of simulation engines), and teleoperation cannot handle (communication latency, occlusion attenuation, multi-vehicle channel congestion). CIDI holds exactly this data.
In his WRC keynote, CIDI CEO Hu Sibo revealed a set of numbers: as of the first half of 2026, CIDI’s autonomous driving fleet exceeds 3,400 vehicles, of which cumulative shipments of autonomous mining trucks have surpassed 1,900 units, covering nearly 40 mines globally, with the largest single-mine regular operation scale reaching 220 units. This means 1,900 heavy machines are running daily in more than 40 mines with extreme operating conditions, continuously generating first-hand data across the full chain of perception, decision-making, and execution.
Hu Sibo made a key statement in his speech: “What autonomous driving accumulates is not just algorithms and mileage, but the cognition forged by 1,900 vehicles running daily across more than 40 mines—and ultimately, the machine’s understanding of the physical world.” The subtext: the moat for heavy-duty embodied intelligence is not model parameters, but real-world data accumulation under extreme operating conditions.
Built on this data, CIDI has constructed a three-layer heavy-duty embodied intelligence technology platform: the bottom layer is a unified data engine aggregating multi-dimensional data from production, environment, and vehicles; the middle layer is a heavy-duty world model, with terminal action models handling millisecond-level real-time decisions and cloud-trained models continuously iterating; the top layer is a cluster decision-making and scheduling system capable of orchestrating over 1,000 agents in a single scenario. More importantly, CIDI hasn’t stopped at “transportation”—70% of mining operations are non-transport tasks. The company is extending its “brain” across the full workflow of drilling, blasting, excavating, and loading: excavation robots have shipped 34 units; open-pit mine charging robots have achieved full-process unmanned operation for hole-finding, charging, and vehicle repositioning; smelting slag-pot carriers transport 75-ton slag pots at temperatures exceeding a thousand degrees Celsius; and underground mining and tunneling robots have completed full-machine manufacturing.
Underpinning all of this is a set of commercial fundamentals that stand out starkly in the embodied intelligence sector: first-half 2026 total revenue of 804 million yuan (approximately $119.6 million), up 97% year-on-year; gross profit of 212 million yuan (approximately $31.5 million), up 204% year-on-year; autonomous driving revenue up 107.3% year-on-year. Against an industry backdrop of heavy spending and cash burn, CIDI is one of the few players to have closed the commercial loop with a clear path to profitability.
The Shovel Sellers: Data-Collection Hardware Erupts
If data is gold, then another camp of players at WRC was selling the shovels. Octopus Dynamics showcased three products—a fisheye headband, an EMG wristband, and an exoskeleton isomorphic data-collection glove—collectively called OctoSense, with the core selling point of achieving the world’s first cross-individual zero-shot generalization of EMG signals. Zibianliang set up a no-robot data-collection demonstration area on-site, with three data-collection hardware products unified under a single data output standard. Daxiao Robotics brought its environmental data-collection solution 2.0, using ACE Ego Kit, ACE Data Engine, and ACE Ego Matrix to build a complete pipeline from collection and automatic annotation to cross-embodiment application. BrainCo’s dexterous manipulation data-collection matrix integrates three data sources: real-machine execution, human demonstration, and simulation generation.
More cutting-edge approaches are also emerging. Fourier Intelligence proposed a “brain-computer data collection” model, building a data system around brain, human, and machine, comparing EEG differences across three states: execution, imagination, and teleoperated robot control. BrainCo demonstrated on-site how brain-computer interfaces connect humans and robots: a human generates action intent, sensors collect task-relevant EEG signals, which are converted into executable commands for the robot control system. KanKan AI introduced an ultra-lightweight 56-gram automotive-grade precision heterogeneous data-collection glasses, targeting passive, massive data collection in real scenarios.
JD.com Enters: Trading Scenarios for Data
Unlike the players above, JD.com appeared at WRC as a global strategic partner, unveiling a robotics strategy spanning three dimensions—supply chain, services, and technology: 10 billion yuan in resource investment, 80 robot bases, an after-sales network covering more than 100 countries, and a two-year plan to collect 10 million hours of real-world scenario data.
JD.com’s logic is straightforward: the robotics industry doesn’t lack companies that can build robots; what it lacks is the “utilities infrastructure”—the water, electricity, and gas—to move robots from exhibition booths into millions of homes and thousands of industries. On the supply chain front, JD.com will invest 10 billion yuan in robotics by 2028, aiming to help 100 brands surpass 1 billion yuan in annual sales. JD.com has already established partnerships with more than 200 robotics brands and numerous core component suppliers through its direct-sales model. On the services front, JD.com has set up 8 major repair centers across China, launched the “120” service standard (1-minute response, 2-hour on-site visit, same-day repair), and plans to expand to 50 cities nationwide over the next three years. Overseas, it launched the JoyRobocare service brand, targeting coverage of more than 100 countries.
But the most critical piece is data. JD.com holds a vast trove of ready-made real-world scenarios: shopping-guide robots in JD MALL, sampling robots in 7Fresh supermarkets, transport and sorting robots in logistics warehouses, packing robots in unmanned pharmacies, and operational robots in front-end fulfillment centers. JD Cloud plans to cumulatively collect over 10 million hours of real-world scenario data within two years, sourced from real operating environments including logistics warehouses, retail stores, and health pharmacies. JD.com is also opening its JoyAI large model and JoyInside intelligent voice platform to partners, with brands including Unitree, Deep Robotics, and Yuanluobo already integrated and deployed.
At WRC’s Developers’ Night, someone proposed a five-layer data progression framework: human demonstrations provide knowledge priors, real-machine execution provides real-world experience, simulation provides exploration efficiency, failure data provides boundary breakthroughs, and industrial scenarios provide sustainable closed loops. Among these five layers, the most scarce and most valuable are the last two—failure data and industrial scenario closed-loop data. Because the first three layers can be scaled through open source, hardware, and simulation, while the last two can only be generated through real, on-the-ground industrial operations.
From this perspective, WRC 2026’s data wars have already stratified: Lightwheel and its peers are competing on data scale breadth, Octopus and its peers on data-collection tool precision, Xinghai Tu and its peers on the balance between real and simulated data, JD.com on scenario coverage density, and CIDI on a track no one else can even enter—closed-loop industrial data under extreme operating conditions.
While more than 300 companies in Yizhuang’s exhibition halls teach robots how to grasp, sort, and tighten screws, 1,900 autonomous mining trucks are operating across 40 mines, in environments of dust, extreme heat, and heavy mixed traffic, continuously producing the industry’s most scarce “dirty data.” The endgame for humanoid robots is entering homes, but before that, a fleet of machines must first descend into hell. And the data that emerges from hell is the one moat that truly cannot be replicated.
Source link
Author

- Ytv Market News
- Share-market news writer and analyst with deep experience covering equities, commodities, forex, and cryptocurrencies for readers in the USA, UK, Canada, and Australia. Ytv Market News delivers timely market updates, practical trading insights, and clear explanations of macro and company-level catalysts that move prices. Combines on-the-ground financial reporting with technical analysis, using concise charts and actionable ideas to help investors and traders make smarter decisions.
Latest entries
AustraliaAugust 25, 2026Peet FY26 earnings: Profit and dividend surge on record sales
Crypto NewsAugust 25, 2026WBT’s new all-time high comes as crypto infrastructure gets more institutional
Politics News TodayAugust 25, 2026Houston police caught on video jumping into Uber to chase a fleeing suspect
Market Movers TodayAugust 25, 2026Okta (OKTA) Q2 Earnings Report Preview: What To Look For
