Open-Source Embodied Intelligence
Introduction: From Demo-Driven Narratives to a Complete Open-Source Ecosystem
In 2025, embodied intelligence reached a critical inflection point, moving from the laboratory into the real physical world. In previous years, the dominant narrative was largely demo-driven: robots were showcased performing smooth, impressive movements in carefully controlled environments. In 2025, however, China’s open-source embodied intelligence ecosystem began to break decisively beyond these constraints.
As universities, research institutions, and technology companies increasingly invested in the field, China’s open-source embodied intelligence ecosystem evolved from the release of individual models into a full-stack ecosystem spanning models, benchmarks, datasets, software stacks, and robot hardware. Open source has not only lowered the barriers to developing embodied intelligence, but has also encouraged the industry to confront the long-tail challenges of the physical world through standardized benchmarks and large-scale, high-quality datasets. This chapter provides a comprehensive overview of the major breakthroughs in China’s open-source embodied intelligence ecosystem in 2025.
1. Open-Source Foundation Models for Embodied Intelligence: From Point Solutions to a Capability Matrix
Over the past year, open-source foundation models for embodied intelligence have moved beyond a single vision-language-action (VLA) architecture. To enable robots to move beyond the lab and perform useful tasks in the real world, research institutions and technology companies have increasingly focused on high-precision spatial perception and genuine physical-world interaction, giving rise to a more diversified and specialized model landscape.
1.1 Spatial Perception Models: Helping Robots “See” the World Clearly
Spatial perception is a prerequisite for robots to interact with the physical world. In early 2026, Ant Group’s LingBot officially open-sourced its next-generation spatial perception model, LingBot-Depth. The model employs a Masked Depth Modeling (MDM) approach, deliberately masking portions of depth information during training and forcing the model to learn how to reconstruct the missing information. Trained on data captured with Orbbec’s Gemini 330-series binocular 3D cameras, LingBot-Depth demonstrates strong performance in challenging scenarios involving transparent and reflective objects. It embodies a broader technical philosophy: using software to compensate for hardware limitations.
1.2 Action Models: Helping Robots “Do the Right Thing”
Action models serve as the execution core of embodied intelligence. Their primary task is to translate visual observations and language instructions into continuous actions that a robot can execute in the physical world.
Over the past year, the field has evolved along two parallel technical paths. The first is the increasingly mature Vision-Language-Action (VLA) model, which continues to dominate the field. The second is the rapidly emerging World-Action Model (WAM) / Video-Action (VA) model, developed in part to address a key weakness of VLA models: they may understand semantics without adequately understanding the underlying physics of the world. China’s open-source community has produced influential work along both trajectories.
1.2.1 VLA Models: From “Moving” to “Doing the Right Thing”
Over the past year, China’s open-source community has made a series of breakthroughs in VLA parameter efficiency, generalization, and real-robot performance, helping move VLAs from merely being able to execute actions to reliably executing the correct actions. On several dimensions, these models have also surpassed leading closed- and open-source baselines, including Physical Intelligence’s π0 series.
Looking back along the open-source timeline, this evolution can be understood as a progression from architectural innovation, to extreme parameter efficiency, and finally to full-stack engineering and deployment at scale.
One of the early milestones was CogACT, led by Microsoft Research Asia in collaboration with Tsinghua University, the University of Science and Technology of China, and the Institute of Microelectronics of the Chinese Academy of Sciences. CogACT represents one of the Chinese open-source community’s notable contributions to VLA architectural innovation.
Early VLA models such as OpenVLA and Octo commonly discretized actions into tokens and directly reused language-model prediction mechanisms, which could result in a loss of control precision. CogACT instead introduced a modular architecture built around a new paradigm of decoupling cognition from action and employing a diffusion-based action expert. This design subsequently influenced a generation of VLA models, including several Chinese models discussed below, and became an important open-source reference point in the field.
This was followed by an academically driven “small but powerful” line of research. Tsinghua University’s Institute for AI Industry Research (AIR), together with Shanghai AI Laboratory, released X-VLA. X-VLA adopts a streamlined flow-matching architecture specifically designed for stable pretraining on heterogeneous datasets. It became the first fully open-source model—with publicly available data, code, and model weights—to complete a 120-minute autonomous clothes-folding task without assistance. Despite having only 0.9 billion parameters, it set new records across five major simulation benchmarks and won the IROS 2025 AGIBOT World Challenge, demonstrating the viability of small-scale models with strong generalization in embodied intelligence.
Running in parallel was RoboBrain 2.0 from the Beijing Academy of Artificial Intelligence (BAAI), a unified foundation model for perception, reasoning, and planning that further enhanced robots’ autonomous decision-making capabilities in complex, long-horizon tasks. Another milestone was XR-1, the first domestic model to pass the national EI Bench standard for embodied intelligence, marking an important step forward for open-source VLAs in terms of compliance and reliability.
While academia was strengthening the “small but powerful” approach, a second trajectory led by technology companies—focused on full-stack engineering and deployment—gained significant momentum from the second half of 2025 into early 2026.
Among the earliest examples was X Square Robot, founded in late 2023, which officially open-sourced its embodied foundation model WALL-OSS in September 2025. WALL-OSS is a general-purpose embodied foundation model with 4.2 billion parameters.
Architecturally, WALL-OSS introduced a shared-attention + expert-routing (FFN) mechanism that maps language, vision, and action into a unified representation space. This design helps mitigate two major challenges in transferring knowledge from vision-language models (VLMs): catastrophic forgetting and modality decoupling.
For training, WALL-OSS adopts a two-stage strategy consisting of an Inspiration Stage followed by an Integration Stage, following a progression from discrete representations to continuous representations and finally joint modeling. This enables the cognitive capabilities of VLMs to be transferred to physical actions with minimal loss.
More importantly, X Square Robot emphasized the completeness of its open-source release. The company released the pretrained model weights, training code, dataset interfaces, and deployment documentation together, allowing developers with roughly RTX 4090-level compute to run the full pipeline from training to deployment. External teams could reportedly adapt the model to third-party robot platforms in as little as one week, compared with the typical one- to two-month adaptation cycle. As CTO Wang Hao put it, the goal was to give the entire industry access to advanced, general-purpose capabilities at the lowest possible cost, effectively building infrastructure for an embodied-intelligence industry that had long been caught in a cycle of overfitting to demonstrations.
Around the same time, GigaBrain-0, an end-to-end VLA foundation model jointly released and open-sourced by GigaVision and the Hubei Humanoid Robot Innovation Center, introduced a distinctive world-model-driven VLA approach.
GigaBrain-0 was the first domestic VLA foundation model to use a world model to generate training data for real-robot generalization. Its core approach relies on the company’s proprietary world-model platform, GigaWorld, to generate large volumes of synthetic data across Sim2Real, Real2Real, novel-view generation, video generation, and human-video transfer. By expanding the diversity of real-robot data by roughly 10×, the system aims to build a comprehensive embodied dataset at a much lower cost than collecting real-robot data alone.
At the model level, GigaBrain-0 supports inputs including images, point clouds, text, and robot proprioceptive states. It incorporates depth information to improve 3D spatial perception and uses subgoal decomposition and end-effector trajectory prediction to enable structured reasoning. These capabilities have translated into strong performance on flexible, long-horizon, mobile manipulation tasks such as folding clothes, organizing toilet paper rolls, clearing tables, and moving boxes.
The team subsequently released GigaBrain-0.1, trained on tens of thousands of hours of data. It surpassed models such as π0.5 in the RoboChallenge real-robot evaluation, demonstrating the scalability of the “world model as a data engine” approach.
Spirit AI, founded in early 2024, open-sourced its embodied foundation model Spirit v1.5 in January 2026. The model topped the RoboChallenge real-robot leaderboard with an overall score of 66.09 and a success rate of 50.33%, becoming the first domestic model to surpass the π0.5 baseline since the benchmark was launched.
Spirit v1.5 uses a unified VLA architecture. Its key innovation lies in its pretraining data strategy: instead of relying on highly curated and tightly controlled “clean data,” it adopts a more diverse, open-ended, and weakly controlled data-collection paradigm. Data collectors are encouraged to act freely around the task objective, allowing the dataset to naturally capture a broad range of atomic skills, including grasping, insertion, organization, bimanual coordination, and error recovery.
Ablation experiments showed that, with the same amount of pretraining data, models trained on more diverse data required roughly 40% fewer iterations to reach the same performance on new tasks. This supports the conclusion that task diversity can be more important than simply increasing the number of demonstrations for a single task. Spirit AI also open-sourced the base-model weights, inference code, and usage examples, providing the research community with an open-source technical path distinct from the π series.
Shortly afterward, Ant Group’s LingBot released LingBot-VLA, explicitly positioning it as a “one brain, many robots” industry foundation model. Rather than targeting a particular robot platform, LingBot-VLA is designed as a general-purpose action foundation model that can be reused across different robot embodiments and transferred across tasks.
To support this generalization, LingBot-VLA was pretrained on more than 20,000 hours of real-robot data covering nine major bimanual robot configurations. One of its key technical features is its integration with the high-precision spatial perception model LingBot-Depth, released around the same time, allowing depth information to be incorporated directly into the action model and further improving task success rates.
On Shanghai Jiao Tong University’s open-source embodied benchmark GM-100, which contains 100 real-world manipulation tasks, LingBot-VLA achieved an average cross-embodiment success rate of 17.3% across three different real-robot platforms, up from 13.0% for π0.5, setting a new record on the real-robot evaluation at the time.
Its positioning as a foundation model is also reflected in the completeness of its open-source release and the attention it received from the developer community. In addition to the model itself, Ant Group open-sourced the post-training code, enabling developers to adapt the general-purpose foundation model to their own robot platforms and tasks at relatively low post-training cost. Its combination of low entry barriers and reusability quickly attracted significant attention from the open-source community, making LingBot-VLA one of the most-starred Chinese open-source embodied action models on GitHub during the period.
Unlike the “general-purpose foundation model” strategies pursued by the companies above, Dexmal has taken a differentiated “embodied-native” approach. In February 2026, Dexmal unveiled DM0 at its inaugural Technology Open Day and fully open-sourced the 2.4B-parameter version, including its code, model, and parameters and inference code for 30 tasks.
Here, “embodied-native” means that the model is trained from scratch and deeply integrates multimodal information from the internet with sensor data specific to embodied environments, rather than being derived from a general-purpose VLM.
DM0 has three core characteristics: pretraining on multi-source data; multi-task, cross-embodiment pretraining; and spatial-reasoning chains of thought. During pretraining, it systematically combines manipulation, navigation, and whole-body control tasks across eight robot platforms with substantially different hardware configurations, enabling strong cross-embodiment generalization. Its spatial reasoning chain connects environmental perception, task understanding, motion planning, and fine-grained execution into a closed loop.
In the RoboChallenge real-robot evaluation, DM0 ranked first in both the single-task and multi-task categories and topped the global leaderboard. The team also open-sourced the modular embodied development framework Dexbotic 2.0, discussed in greater detail in Chapter 4.
Taken together, these efforts demonstrate that China’s open-source VLA ecosystem in 2025–2026 is no longer satisfied with simply releasing a model that can run. Instead, the emerging standard is a fully reproducible full-stack solution, publicly verifiable real-robot benchmark results, and validation in real-world production environments. The ecosystem is expanding from general-purpose foundation models toward embodied-native models, and from individual models toward a broader full-stack capability matrix—helping move embodied foundation models toward genuinely open and widely accessible technology.
1.2.2 WAM/VA: Helping Robots “Imagine First, Then Act”
If VLA models are primarily designed to understand instructions and map them to actions, the most notable paradigm shift of 2025 came from World-Action Models (WAMs), also known as Video-Action (VA) models.
The core idea is to use a pretrained video diffusion model as the backbone and jointly model future visual states and robot actions within a unified framework. This allows the model to inherit physical and spatiotemporal priors embedded in massive amounts of internet video, giving robots a form of predictive intelligence: they can first imagine how the world will look after an action and then infer what they should do now.
This approach directly addresses a fundamental weakness of VLA models: while VLAs can provide strong semantic priors, they often lack sufficiently rich priors about physical dynamics. Although WAM/VA emerged later than VLA, its potential advantages in generalization and data efficiency have made it one of the most promising frontiers in embodied foundation models, with China’s open-source community playing a particularly active role.
LingBot-VA, released by Ant Group’s LingBot in early 2026, is a representative open-source effort along this trajectory. It has been described as the world’s first autoregressive video-action world model. While predicting the next visual state of the world, the model simultaneously generates the action commands required for the robot to bring about that predicted state. In this sense, the robot can “reason and act simultaneously,” translating the predictive capabilities of a world model into physical action.
On real-robot tasks, LingBot-VA reportedly improved success rates by approximately 20% on average compared with strong industry baselines such as π0.5. In simulation, it exceeded 90% success on RoboTwin 2.0 for the first time and achieved an average success rate of 98.5% on LIBERO, setting new records at the time. Its model weights and inference code have since been fully open-sourced.
Motus, jointly developed by Generative Intelligence and Tsinghua University, is another representative Chinese open-source project in the WAM/VA space. Generative Intelligence officially released and open-sourced the general-purpose foundation world-action model in February 2026. Built on the team’s original UniDiffuser unified modeling framework, Motus brings language, video, and action into a single framework. A single training process supports multiple capabilities, including VLA, video generation, inverse dynamics, and joint video-action prediction. The code, paper, and model weights have all been fully open-sourced.
An international counterpart is NVIDIA’s DreamZero, a 14-billion-parameter WAM built on a pretrained video-diffusion backbone, Wan2.1-I2V-14B. DreamZero shares the same core concept as the models described above, further indicating that WAM/VA is emerging as a potential next-generation paradigm following VLA.
Notably, China’s open-source community has already established multiple strong efforts in this emerging area, including LingBot-VA and Motus. These projects are competing at the international frontier along both architectural innovation and engineering efficiency, giving China’s open-source ecosystem a significant presence in this rapidly developing field.
1.3 World Models: Interactive World Modeling
It is important to distinguish between two technical approaches that are both commonly described as “world models” in this report.
The first is the control-loop-oriented approach discussed in the previous section: WAM/VA models generate future visual states while directly outputting robot actions.
The second, which is the focus of this section, is interactive world modeling. Its core capability is to use a user’s control signals—such as a keyboard, mouse, game controller, or physical actions—as conditioning inputs to generate a virtual world in real time that is explorable, interactive, and temporally consistent.
Unlike traditional simulators, which rely on rendering engines to construct and render scenes frame by frame, generative world models directly imagine how the environment should change in response to an interaction. Given the current visual state and an interaction command, the model continuously generates a video stream in which an agent can effectively “exist” and interact with the generated environment.
With Google Genie 3 as a major reference point, this direction saw a wave of breakthroughs from Chinese open-source teams in the second half of 2025 and early 2026.
The motivation is straightforward: collecting complex, long-horizon real-robot data is extremely expensive and inherently difficult to scale. A sufficiently realistic, interactive, and generalizable generative world could provide virtually unlimited and reproducible interaction and training data for embodied intelligence, autonomous driving, games, film and television production, and other fields.
Kunlun Tech’s Skywork AI was an early open-source pioneer in this direction, beginning its work even before Google released Genie 3. In May 2025, it open-sourced Matrix-Game, a 17-billion-parameter model described as the industry’s first open-source spatial-intelligence foundation model with more than 10 billion parameters.
Starting from a single image, Matrix-Game is an interactive world foundation model designed around game-world modeling. Users can freely explore generated Minecraft environments through keyboard controls such as W/A/S/D, jumping, and attacking, while using the mouse to control the camera view.
Matrix-Game also introduced GameWorld Score, an evaluation framework specifically designed for interactive world generation. It measures model capabilities across four dimensions: visual quality, temporal consistency, action controllability, and understanding of physical rules.
Kunlun Tech followed this with the open-source release of Matrix-Game 2.0 and Matrix-3D during its Technology Release Week in August 2025.
Matrix-Game 2.0 was described as the industry’s first open-source world model capable of real-time, long-sequence interactive generation in general-purpose environments. It can generate continuous interactive video at 25 FPS for minutes at a time across urban, wilderness, and other environments and across different visual styles.
Matrix-3D, meanwhile, starts from a single image to generate trajectory-consistent panoramic video and reconstruct a navigable 3D environment, targeting capabilities similar to those demonstrated by Fei-Fei Li’s World Labs.
Following Kunlun Tech, Tencent Hunyuan also made sustained investments in interactive world modeling, becoming another major player through a strategy of continuous iteration and end-to-end open sourcing.
After open-sourcing HunyuanWorld-1.0, the first 3D world-generation model from Tencent to support physics simulation, in July 2025, Tencent released World Model 1.1 (WorldMirror) in October with support for multi-view and video inputs. In December 2025, Tencent Hunyuan released and open-sourced Hunyuan World Model 1.5 (HY WorldPlay), designed specifically for real-time interaction.
Unlike earlier systems that relied primarily on offline generation, WorldPlay is built around an autoregressive diffusion model. Users can create an interactive world from text or images and control the movement and orientation of the virtual camera in real time using a keyboard, mouse, or game controller. Generation can reach 24 frames per second.
The system supports first- and third-person perspectives, event-triggered scene changes such as smoke and explosions, and 3D scene reconstruction.
To address the tension between real-time generation and long-term consistency, Tencent introduced a reconstructed-context memory mechanism, which dynamically reconstructs information from previous frames to maintain long-term geometric consistency, as well as WorldCompass, a reinforcement-learning-based post-training framework specifically designed for long-sequence autoregressive video models.
Tencent describes WorldCompass as a comprehensive world-model framework covering the entire pipeline from data and training to streaming inference and deployment. On its benchmarks, WorldPlay reportedly outperformed all compared models on visual quality and long-term geometric consistency, while lagging behind some individual models slightly in camera-rotation accuracy. HY WorldPlay is currently open-sourced on GitHub and Hugging Face.
Entering 2026, Ant Group’s open-source LingBot-World emerged as one of the most closely watched open-source projects in this area. Released in January 2026, it was described by multiple media outlets as the first open-source world model capable of competing with Google Genie 3.
At its core, LingBot-World-Base builds on video-generation technology and is powered by a Scalable Data Engine. By learning physical laws and causal relationships from large-scale game environments, the model enables real-time interaction with the worlds it generates.
LingBot-World demonstrated strong performance across several key dimensions. In terms of long-horizon temporal consistency, multi-stage training and parallel acceleration enabled nearly 10 minutes of continuous, stable generation without significant degradation. Even after the camera moved away from a scene for up to 60 seconds and returned, core objects retained consistent geometry and appearance.
For real-time interaction, the system can generate at approximately 16 FPS, with end-to-end interaction latency kept below one second. Users can control characters and camera viewpoints in real time using a keyboard and mouse, and can even use text commands to trigger environmental changes such as weather and visual style.
Its model weights, inference code, and technical report have all been fully open-sourced.
Notably, the open-source release of LingBot-World occurred almost simultaneously with the public launch of Google Project Genie, the experience platform built around Genie 3. Whereas Google made the experience available only to subscribers, LingBot-World released its weights and code openly, giving developers relatively low-barrier access to an industrial-scale interactive world model for the first time.
From Kunlun Tech’s Matrix series, to Tencent Hunyuan’s WorldPlay, and finally to LingBot-World, interactive world modeling has clearly become a major frontier for China’s open-source AI community. Across these projects, two challenges consistently emerge as the primary engineering targets: long-term consistency through spatial memory, and real-time interaction through low latency and high frame rates.
2. Evaluating Open-Source Embodied Models: Breaking the Illusion of “Visual Realism”
As model capabilities continue to advance, traditional evaluation frameworks are increasingly unable to capture how well embodied AI systems actually perform in complex environments. In 2025, Chinese researchers and industry players joined forces to develop a two-dimensional evaluation framework spanning both simulation and real-robot testing, specifically designed to assess action models under increasingly realistic conditions.
2.1 Simulation Evaluation: From Basic Skills to Complex Interaction
Simulation evaluation is the first checkpoint for validating model capabilities. Compared with expensive and difficult-to-scale real-robot testing, simulation environments provide large-scale evaluation at a fraction of the cost, with reproducible conditions and controllable variables. They have therefore become essential infrastructure for iterative policy development.
In addition to widely used international benchmarks such as LIBERO for lifelong learning, CALVIN for long-horizon language-conditioned manipulation, and SimplerEnv for real-to-sim evaluation, China’s open-source community showed a clear trend in 2025: moving from basic skill evaluation toward complex interaction and higher-level reasoning. This shift has led to the development of a number of simulation benchmarks that have gained broad adoption in the international research community.
RoboTwin 2.0, designed for bimanual manipulation, is one of the most representative and influential examples. The project was jointly developed by Shanghai Jiao Tong University, the University of Hong Kong, Shanghai AI Laboratory, and other institutions. Its early version won the Best Paper Award at the ECCV 2024 Workshop, while version 1.0 was selected as a CVPR 2025 Highlight.
Unlike conventional static task suites, RoboTwin 2.0 is essentially an integrated framework combining a scalable data generator with a unified evaluation benchmark. It includes the RoboTwin-OD object library, covering 147 object categories and 731 objects annotated with semantic and manipulation information. With the help of multimodal large language models (MLLMs) and closed-loop simulation-in-the-loop feedback, the system can automatically synthesize task programs covering 50 bimanual manipulation tasks across five robot embodiments.
RoboTwin has now fully open-sourced its data generator, benchmarks, datasets, and code. It also served as the official platform for the CVPR 2025 Bimanual Manipulation Challenge, making it an important piece of shared infrastructure for bimanual manipulation research in China.
For evaluating VLA models and higher-order reasoning, VLABench, developed by Fudan University, fills an important gap. As the first robot manipulation benchmark specifically designed for VLA models, with language-based instructions and long-horizon reasoning tasks, VLABench goes beyond measuring whether a policy can execute individual actions correctly.
Instead, it evaluates multimodal reasoning and zero-shot task-planning capabilities across dimensions including vision, language, planning, and common-sense reasoning. In doing so, it pushes embodied-model evaluation beyond simply asking whether a robot can “perform the right action” toward whether it can “understand the logic of the task.”
At the same time, simulation evaluation is evolving from team-specific task suites toward shared platforms and evaluation-as-a-service.
During its Embodied Intelligence Open-Source Week in September 2025, Shanghai AI Laboratory, leveraging its full-stack embodied intelligence engine Intern-Robotics, introduced an evaluation platform for multimodal navigation and manipulation in high-fidelity environments. The platform provides the community with open-source evaluation tools, baseline methods, datasets, and evaluation services.
Its navigation benchmark focuses on vision-language navigation in physically realistic environments, while its manipulation benchmark emphasizes long-horizon instruction following with reasoning. Building on this infrastructure, the IROS 2025 Challenge was opened to researchers worldwide, with evaluation services made available to the community on an ongoing basis.
In 2026, Shanghai AI Laboratory further released EBench, a systematic simulation benchmark for embodied manipulation. EBench includes 26 task types and labels tasks across five dimensions—scene, atomic skill, duration, precision, and mobility—while providing 794 test tasks for fine-grained capability diagnosis and generalization evaluation.
Together, these developments have further strengthened the role of simulation benchmarks as community-built, fair, and reproducible infrastructure.
Overall, from bimanual coordination and long-horizon compositional generalization to higher-level language reasoning, China’s open-source simulation benchmarks are providing increasingly low-cost yet increasingly realistic environments for both model pretraining and capability diagnosis.
2.2 Real-Robot Evaluation: The Ultimate Test in the Physical World
No matter how realistic a simulation may be, real-robot evaluation remains the ultimate test of embodied intelligence.
For years, real-robot evaluation has faced several fundamental challenges: poor reproducibility, a lack of standardized evaluation protocols, and high testing costs. As a result, the various demo-driven narratives promoted by different companies have been difficult to compare systematically and even harder to independently reproduce.
As Turing Award winner Andrew Chi-Chih Yao has emphasized, the embodied AI industry urgently needs to move from everyone speaking their own language toward standardized evaluation. From 2025 to 2026, the emergence of two major open-source real-robot evaluation platforms has begun to address this longstanding problem in a systematic way.
RoboChallenge, jointly initiated by Dexmal and Hugging Face, is described as the world’s first large-scale real-robot evaluation platform for embodied intelligence. It aims to create an open, fair, and highly reproducible “real-world testbed.”
Since its launch on October 15, 2025, the platform has deployed a 20-robot cluster covering four major robot platforms, including UR5, Franka Panda, ARX5, and ALOHA, establishing a stable and diverse remote physical-testing network.
It has also open-sourced the Table30 dataset, consisting of 30 standardized tabletop tasks across nine major task categories.
One particularly notable aspect is its open-community governance model. In November 2025, Dexmal and several other organizations jointly established the RoboChallenge organizing committee, helping shift evaluation from isolated experiments toward community-driven consensus and shared infrastructure. This marked an important step toward standardizing real-robot evaluation.
The platform’s first annual report, released on January 30, 2026 (read the full report here), was based on tens of thousands of rigorous remote real-robot tests conducted over several months, covering Q4 2025 through Q1 2026. The report provides an empirical picture of the current capability limits of VLA models.
Its findings are particularly revealing. Basic tasks such as stacking bowls and moving objects into boxes have approached maturity and are increasingly becoming the embodied-AI equivalent of “Hello World”. More complex tasks involving multi-step sequential decision-making, long-horizon planning, and fine-grained dexterous manipulation—such as organizing paper cups and making sandwiches—have continued to show low success rates, with some tasks approaching zero.
Even the top-performing model on Table30 achieved an overall success rate of only around 50%, while success rates on fine-grained manipulation tasks remained below 15%.
This publicly accumulated collection of real-world failures effectively serves as a shared failure set and empirical yardstick, providing the broader research community with valuable data for model iteration and exposing where current systems still struggle.
GM-100, released in early 2026 by the team led by Yonglu Li at Shanghai Jiao Tong University, takes its name from “Great March,” or the Long March, reflecting the team’s view that building a comprehensive real-world benchmark is a long-term, labor-intensive undertaking.
The benchmark contains 100 tasks, with approximately 100 training trajectories and 30 test trajectories per task, for a total of roughly 13,000 real-world manipulation trajectories.
Its design is rooted in a data-centric approach to embodied intelligence. The team found that existing datasets remain heavily concentrated around three basic categories—pick, hold, and place. GM-100 deliberately takes the opposite approach by focusing on long-tail and fine-grained manipulation tasks, such as threading candied fruit onto a skewer, opening drawers, pressing a desk-lamp switch, and organizing small objects.
These tasks were constructed through a pipeline of human-interaction primitive analysis → LLM-generated candidates → expert screening and refinement.
The resulting benchmark deliberately captures a number of counterintuitive failure modes. Tasks that humans find difficult may sometimes be relatively easy for robots, while seemingly trivial human actions can fail repeatedly because of factors such as robot morphology, object material, object placement, or instruction interpretation.
GM-100 also goes beyond the conventional success rate (SR). It introduces partial success rate (PSR) and action prediction error.
PSR makes it possible to quantify how much of a multi-step task has actually been completed, while action prediction error measures how accurately a model imitates previously unseen trajectories. Together, these metrics are intended to discourage models from taking shortcuts or gaming the leaderboard, and instead encourage researchers to focus on genuine generalization and imitation capabilities.
The team has already demonstrated the benchmark’s ability to differentiate among several mainstream models, including Diffusion Policy, π0, π0.5, and GR00T.
Equally important is GM-100’s philosophy of community-driven development. Rather than positioning itself as an authoritative judge, the team describes its role as providing the infrastructure and stage for the community.
It has open-sourced detailed specifications for all 100 tasks, a complete materials list—including links to specific products on Taobao—and roughly 130 real-world manipulation trajectories for each task. Open-source models that successfully pass validation can also receive a “verified” label.
This approach resembles the decentralized, mechanism-driven evaluation model of LMArena in the LLM community, lowering the barrier to reproduction and encouraging broader participation.
According to the team, GM-100 is expected to gradually expand its task library to 300 and eventually 1,000 tasks, while also introducing cross-platform evaluation across different robot embodiments.
3. Open-Source Datasets: Building the Data Pyramid for Embodied Intelligence
3.1 Datasets for Spatial Perception
As spatial intelligence continues to advance, China’s open-source community has released a large number of perception datasets containing depth information, point-cloud data, and 3D semantic annotations. These datasets provide a rich source of training data for enabling robots to understand complex three-dimensional environments.
One of the most representative examples is LingBot-Depth-Dataset, open-sourced by Ant Group’s LingBot in March 2026. It is currently the largest open-source RGB-D dataset for real-world scenes, containing 3 million high-quality sample pairs—2 million collected from real-world environments and 1 million generated through rendering. The dataset totals 2.71 TB and covers six mainstream depth-camera models.
Each sample provides an RGB image, the sensor’s raw depth map, and a ground-truth depth map, making the dataset directly applicable to both depth estimation and depth completion training and evaluation. It fills an important gap in publicly available spatial-perception data captured from real-world environments.
Beyond purely visual depth perception, vision-tactile fusion has also emerged as a major area of interest in spatial-perception datasets.
OmniViTac, jointly released by Itsa Stone and the National University of Singapore along with four other institutions, is the first large-scale cross-embodiment vision-tactile-action alignment dataset. It has demonstrated strong performance on contact-rich manipulation tasks.
Meanwhile, Baihu-VTouch, open-sourced by the National-Local Joint Innovation Center for Humanoid Robots together with Weitai Robotics, is the world’s largest cross-embodiment vision-tactile multimodal dataset. It provides an important foundation for enabling robots to develop a more accurate understanding of the physical world through touch.
3.2 The Data Pyramid for Embodied Action Models
In the field of action models, the industry is gradually developing a clear “data pyramid” structure, with representative open-source projects emerging at each level:
| Tier | Data Type | Scale | Representative Open-Source Projects and Tools |
|---|---|---|---|
| Tier 1 | High-fidelity real-robot teleoperation data | Hundreds of thousands of hours | RoboMIND: Released by the National-Local Joint Innovation Center for Embodied Intelligent Robots and other organizations. Version 1.0 contains 107,000 real-robot trajectories spanning 4 embodiments, 479 tasks, and 38 skills. Version 2.0 expands the dataset to more than 310,000 trajectories, 6 robot embodiments, and 739 tasks, and adds 12,000 tactile-enabled trajectories. Global downloads have surpassed 6 million. AgiBot World: Released by AgiBot and partners, with data collected in a 4,000 m² dedicated data-collection facility. It contains more than 1 million real-robot trajectories covering 100+ robots, 5 major environments, and 1,000+ tasks, making it one of the largest open-source real-robot datasets in the world. Galaxea Open-World Dataset: Released by Galaxea, this dataset was collected using homogeneous R1 Lite robot platforms and contains 500 hours of real-world mobile manipulation data across diverse environments, including homes, kitchens, and retail settings. RoboCOIN: Released by the Beijing Academy of Artificial Intelligence and partners, this dataset covers 15 heterogeneous robot platforms and contains more than 180,000 demonstration trajectories spanning 421 tasks and 16 types of environments. It is currently one of the most extensive bimanual real-robot datasets in terms of robot embodiments and provides highly detailed annotations. LET: Released by Leju Robotics, this dataset was collected using full-size humanoid robots from the Kuavo series. The first open-source release contains more than 60,000 minutes of real-robot data, covering 31 tasks and 117 atomic skills, making it one of China’s largest humanoid real-robot datasets. Ruiyuan Real-Robot Dataset: Released by RealMan Intelligent Robotics, this dataset was collected across 10 real-world scenarios at the Beijing Humanoid Robot Data Training Center. It achieves 100% modality completeness and is positioned as the world’s first high-quality real-robot dataset with comprehensive multimodal coverage. |
| Tier 2 | Low-fidelity passively collected human data | Millions of hours | Representative collection methods include UMI handheld grippers, wearable exoskeletons, and first-person (egocentric) recording devices. These approaches enable low-cost, large-scale collection of human manipulation trajectories without relying on real-robot teleoperation. HORA: Released by Shuto Technology, HORA is described as the industry’s first embodied multimodal dataset extracted from human videos captured in real-world environments. It contains more than 150,000 high-quality trajectories and establishes an end-to-end data pipeline from human demonstration videos to robotic learning for the first time. EgoLive: Released by JD.com, this dataset was collected using custom head-mounted devices and contains 1,680 hours of stereo video at 60 FPS and 2160p, along with more than 65,000 manipulation segments covering 346 real-world tasks. It is currently the largest open-source egocentric interaction dataset of its kind. TASTE-Rob: Released by The Chinese University of Hong Kong, Shenzhen, this dataset contains 100,856 egocentric videos of human–object interactions precisely aligned with language instructions. It is the first large-scale HOI dataset designed for generalizable robotic manipulation (CVPR 2025). |
| Tier 3 | Internet video data | Tens of millions of hours | Large-scale video platforms and open-source video datasets provide vast amounts of human motion and interaction priors for pretraining world models and VLA models. This data is abundant and can be collected at near-zero cost, but generally lacks precise action annotations. |
| Tier 4 | Simulation data | Long-tail coverage | High-fidelity simulation platforms such as Genie Sim 3.0 can generate synthetic data covering long-tail scenarios and hazardous tasks that are difficult or costly to collect in the real world. InternData-A1: Released by the Shanghai Artificial Intelligence Laboratory, this dataset contains more than 630,000 simulated trajectories totaling over 7,400 hours. It covers multiple robot morphologies and complex interaction scenarios and has been directly adopted by several major foundation models. ArtVIP: Released by the Beijing Innovation Center for Humanoid Robotics, ArtVIP provides 206 high-fidelity articulated digital assets across 26 categories, reducing the cost of virtual debugging by up to 80%. AgiBot Digital World Dataset: Released by AgiBot, this dataset is automatically generated using the company’s proprietary large-scale simulation framework. It covers five major environment categories—homes, retail stores, offices, restaurants, and industrial settings—as well as 180+ object categories, 9 material types, and 12 core skills. DexGraspNet: Released by Professor He Wang’s research group at Peking University, DexGraspNet is a simulation dataset for dexterous-hand grasping. Version 1.0 contains 1.32 million grasp configurations involving 5,355 objects across 133 categories. Version 3.0 further expands the dataset to 17 million validated grasp poses covering more than 174,000 objects. |
4. Open-Source Software Stack: End-to-End Infrastructure from Training to Deployment
In 2025, competition in open-source embodied intelligence expanded beyond models themselves to encompass the entire upstream and downstream software stack. Following the end-to-end workflow of how a model is created from data, runs on a system, and is ultimately deployed on a robot, China’s open-source community has built a layered software infrastructure spanning model-training toolchains, robot operating systems and motion control, simulation platforms, and edge deployment. Together, these efforts have produced a growing number of projects with global visibility and influence.
4.1 Model Training Toolchains: From Data Generation to Model Iteration
Training an embodied model is an end-to-end pipeline comprising three successive stages: data collection and teleoperation → model development and fine-tuning → training and post-training. Representative open-source tools have emerged across each stage within China’s open-source community.
At the front end of the pipeline, data collection and teleoperation provide the primary entry point for generating high-quality real-robot data. OpenWBT (Open Whole-Body Teleoperation), open-sourced by Galbot in collaboration with the Tsinghua University Yili Lab, is a representative example. Built on R2S2 technology, it enables whole-body teleoperation of humanoid robots such as Unitree G1 and H1 using Apple Vision Pro and handheld controllers. It closes the loop between virtual simulation and real-robot data collection while incorporating atomic skill reuse, and is fully open-sourced under the Apache 2.0 license. Tools of this kind directly support the production of the real-robot data at the top of the “data pyramid” described in Chapter 3 and serve as the starting point for the entire toolchain.
Moving to VLA model development and fine-tuning, the choice of development framework has a direct impact on how efficiently researchers can reproduce, compare, and iterate on models. Dexbotic, open-sourced by ForceRobo, is a PyTorch-based, all-in-one development toolkit for VLA models. It covers the full pretraining–fine-tuning–inference–evaluation lifecycle through a unified data format (Dexdata), its proprietary foundation model DexboticVLM, and built-in pretrained models including π0, CogACT, and OFT. Its experiment-centric development paradigm allows researchers to quickly reproduce and compare mainstream approaches. Across tests on five major simulation platforms, Dexbotic improved the performance of conventional VLA approaches by up to 46.2%, while achieving a 100% success rate on a real-robot stacking task.
If the tools above address how to efficiently develop a VLA model, starVLA, jointly open-sourced by a team from the Hong Kong University of Science and Technology and the broader open-source community, tackles a more fundamental infrastructure problem: how to evaluate the growing number of VLA approaches under fair, transparent, and reproducible conditions. In response to what the project describes as a “Tower of Babel” problem in the current VLA landscape—fragmented architectures, tightly coupled pipelines, and inconsistent evaluation standards—starVLA introduces a Lego-style modular architecture built around a Backbone–Action Head design. It decouples the training infrastructure, pluggable foundation-model backbones, and action experts into interchangeable building blocks, allowing researchers to replace an action head or backbone by changing just a single line of configuration. More importantly from a theoretical perspective, the authors introduce the concept of Generalized VLA, providing a unified conceptual framework for systematic research in the field. Guided by an engineering philosophy of “staying focused and avoiding unnecessary reinvention,” starVLA has been described as an “iPhone moment for embodied intelligence” and has surpassed 2.9k GitHub stars, making it one of the most prominent open-source projects of its kind in China.
It is important to note that these Chinese open-source toolchains have not developed in isolation. They are deeply embedded in the global open-source ecosystem, with Hugging Face’s LeRobot serving as one of the most important common foundations. As one of the most prominent open-source embodied-intelligence projects globally and a de facto general-purpose development foundation for embodied AI, LeRobot is built on PyTorch and integrates models, datasets, and tools into a unified framework, substantially lowering the barrier to entry for embodied AI development. Its ecosystem is further strengthened by low-cost open-source robot arms such as the SO-100 and SO-101 and the Hugging Face Hub community. As of the v0.5.0 release in early 2026, the project had expanded to support whole-body control of humanoid robots. For China’s open-source community, LeRobot is particularly important as a globally recognized reference point and collaboration platform. On the one hand, Chinese open-source projects can gain broader global visibility by adopting and integrating with LeRobot standards. On the other, Chinese developers are actively contributing code and models back to the international ecosystem. For example, Unitree Robotics developed the unitree_lerobot training framework for its robot platforms, enabling humanoid robots such as the G1 to receive full support in the LeRobot ecosystem. Chinese models such as Tsinghua AIR’s X-VLA have also been officially integrated, while LeRobot’s Chinese-language tutorials written by Zihau Ge were incorporated into the official Hugging Face documentation ecosystem (see Chapter 6). In this sense, China’s open-source embodied-AI community is increasingly becoming part of the global development mainstream through participation, adaptation, and reciprocal contribution.
At the end of the pipeline, training and post-training engines are emerging as a third major scaling pathway after data and model architectures. Reinforcement learning (RL) is becoming increasingly important in this role. Although the ecosystem of general-purpose RL training frameworks has expanded rapidly, with projects such as verl, AReaL, slime, and TRL, most are designed for reasoning-oriented large language models—the “brains” of AI systems. Embodied intelligence, however, has a distinctive render–train–infer integrated workflow: models must interact repeatedly with GPU-accelerated physics simulators, creating intense competition for both compute and GPU memory. As a result, general-purpose frameworks are often poorly suited to the requirements of embodied training. Against this backdrop, RLinf, jointly open-sourced by Tsinghua University, the Beijing Zhongguancun Academy, UnknowAI, and multiple partner institutions, was developed to address the gap in large-scale RL training systems for embodied intelligence. Technically, RLinf introduces its proprietary M2Flow mechanism (Macro-to-Micro Flow) and supports three execution modes—shared, separated, and hybrid—within a unified codebase. It can deliver more than 120% training speedups compared with mainstream frameworks, while supporting both embodied “brain” and “cerebellum” systems and remaining compatible with mainstream models such as OpenVLA, OpenVLA-OFT, π0, and LingBot-VLA.
RLinf’s broader significance, however, lies in its open-source impact and ecosystem value. Developed through close collaboration among industry, academia, and research institutions, the project provides fully open-source code, model weights, and systematic documentation, substantially lowering the barrier to embodied RL research and providing a common experimental foundation for exploring scaling laws for embodied RL. At the same time, it offers a lightweight, out-of-the-box path for users with limited computing resources or limited prior experience. This positioning has helped RLinf attract nearly 4,000 GitHub stars, rapidly establishing it as one of the most closely watched open-source infrastructure projects in embodied reinforcement learning, both in China and internationally.
4.2 Robot Operating Systems, Middleware, and Motion Control
If the training toolchain determines how capable a model can become, robot operating systems and middleware determine whether that model can run reliably on a real robot. Alongside the internationally dominant ROS 2 ecosystem, China’s open-source community is accelerating efforts to build an independently controllable foundational software ecosystem for robotics.
In March 2026, AgiBot officially open-sourced its proprietary robot operating system, Lingqu OS (Alpha). Built around the full-size Expedition A2 platform and informed by production-scale deployment experience, the system’s key component is its unified communications middleware framework, AimRT. It supports both Protobuf and ROS 2 Message formats, maintains compatibility with the native ROS 2 ecosystem, and provides an integrated toolchain for bipedal motion-control simulation, training, and deployment.
Meanwhile, OpenLoong, hosted by the OpenAtom Foundation, continues to evolve. Its open-source software stack includes an embodied-intelligence operating system and a whole-body dynamics control framework. The latter adopts a layered, decoupled architecture, supporting deployment across different hardware platforms and scheduling across different middleware systems, thereby providing foundational software services for dexterous manipulation and robust locomotion in humanoid robots. M-Robots OS, an open-source OpenHarmony-based robot operating system, is likewise working to address heterogeneous-hardware compatibility while improving system real-time performance and security. AGIROS, an intelligent robot operating system initiated by the Institute of Software, Chinese Academy of Sciences, is the country’s first open-source community centered on an independently controllable intelligent robot operating system. It brings together more than 60 leading companies, universities, and research institutions under a co-building, co-sharing, and co-governance model. The project has released four versions and includes more than 1,500 foundational packages, supports multiple CPU architectures and a broad range of robot types, and is fully interface-compatible with ROS 2. By integrating the kernel, middleware, and AI layers into a unified full-stack solution, AGIROS is helping drive the large-scale development of China’s domestic robotics software ecosystem.
At the robotics middleware layer connecting high-level applications with low-level hardware, the dataflow-oriented Dora-RS (Dataflow-Oriented Robotic Architecture) has attracted considerable attention in recent years as a low-latency robotics middleware framework. Unlike traditional architectures centered primarily on topic-based publish/subscribe mechanisms, Dora-RS models complex robotic applications as directed graphs consisting of nodes and data flows. Its underlying implementation is written in Rust and emphasizes zero-copy message transmission. It also supports multiple programming languages, including Rust and Python, as well as multiple platforms and distributed deployment. This significantly reduces the performance overhead associated with cross-language and inter-process communication. By simplifying the development of AI-powered robotic applications while improving real-time performance and scalability, Dora-RS has become an important foundational component for engineering embodied intelligence systems. It has also been adapted and demonstrated on domestic projects and platforms, including OpenLoong and the Qinglong robot hardware platform.
Motion control, particularly whole-body control (WBC), is the critical link between algorithmic decisions and physical execution and has also emerged as a relatively independent open-source domain. OpenLoong’s “Qinglong Motion Control Framework,” Beijing Humanoid’s “Embodied Tiangong” motion-control framework, and Unitree Robotics’ open-source unitree_rl_gym reinforcement-learning control environment collectively provide reusable low-level motion-control infrastructure for developers. unitree_rl_gym supports training and deployment of platforms including Go2, H1, H1_2, and G1 in Isaac Gym and MuJoCo. For highly dynamic and robust whole-body control of humanoid robots, Project Instinct, open-sourced by Tsinghua University’s Institute for Interdisciplinary Information Sciences in January 2026, extends the open-source ecosystem into extreme-motion-control scenarios. The project provides an “instinct-level” whole-body control framework spanning algorithms, environments, data planning, and deployment. It has demonstrated humanoid robots performing challenging skills such as parkour and off-road hiking on irregular terrain, connecting rigid-body physics simulation, high-dimensional perception processing, and end-to-end reinforcement-learning deployment in a single pipeline. Its full source code and core toolkits, including InstinctLab, have been released to the community.
4.3 Simulation Platforms: LLM-Driven Digital Twins
Simulation platforms are essential for addressing the scarcity of embodied-AI data and the high cost of trial and error. They serve both as the “data factory” of the training toolchain and as the “rehearsal ground” for models before deployment. In early 2026, AgiBot released Genie Sim 3.0, an open-source simulation platform driven by large language models. Built on NVIDIA Isaac Sim, Genie Sim 3.0 combines 3D reconstruction and visual-generation technologies to create high-fidelity, digital-twin-style environments. Its most notable innovation is its LLM-driven simulation-generation mechanism: developers can enter natural-language instructions and generate simulation scenarios at the scale of tens of thousands within minutes. AgiBot has adopted a comprehensive open-source strategy, releasing the platform code, simulation assets, evaluation tools, and data resources to the public. This sends a clear signal that simulation infrastructure should be developed as an open ecosystem rather than built behind closed doors.
4.4 Developer Platforms and Edge Computing
The final destination of the software stack is deployment—running models efficiently on the robot itself. D-Robotics has bridged the gap between low-level hardware and high-level algorithms through its RDK (Robot Developer Kit) family of development kits and integrated developer platform. Its RDK X5 delivers 10 TOPS of edge inference performance and an 8-core ARM A55 processor, specifically designed for robotics developers. More importantly, through capabilities such as its integration with Volcano Engine’s edge AI foundation-model gateway, the RDK X5 connects the cloud–edge–device pipeline, allowing developers to invoke cloud-based foundation models directly through standard ROS interfaces. This hardware + algorithms + community model significantly lowers the barrier to entry for small teams, makers, and individual developers, accelerating the integration and deployment of diverse intelligent-robot applications.
5. Open-Source Robot Platforms: A Thriving Hardware–Software Ecosystem
Open-sourcing and standardizing robot hardware platforms is the physical foundation for bringing embodied intelligence into large-scale applications. In 2025, domestic robot-platform manufacturers made significant progress in commercialization while simultaneously embracing the open-source ecosystem, forming a complete hardware–software synergy with the open-source software stack described in the previous chapter.
In the area of full-size humanoid reference platforms, the OpenLoong (Qinglong) community, incubated by the OpenAtom Foundation, fully open-sourced the hardware designs, core components, and drive solutions for the Qinglong full-size general-purpose humanoid reference platform. The initiative aims to lower the barriers to developing full-size humanoid robots while fostering a broader industry ecosystem. The National-Local Joint Innovation Center for Embodied Intelligent Robots has likewise continued to advance its Tiangong Open-Source Initiative. In the Embodied Tiangong 3.0 release in 2026, it opened up a comprehensive set of core components, including the Tiangong robot platform, motion-control framework, world model, foundation model and training toolchain, and datasets.
For community-driven, fully open-source prototype platforms, Shanghai-based RoboParty’s roboto_origin represents another approach centered on extreme openness. Designed for developers and enthusiasts, it is a fully open-source, grassroots bipedal humanoid prototype that provides complete access to its mechanical design, electrical architecture, training pipeline, and deployment code. The robot can be assembled using components sourced through general-purpose supply chains. The project has accumulated more than 1,300 GitHub stars and attracted a community of over 1,500 developers, substantially lowering the barrier to entry for humanoid-robot hardware development and helping broaden the adoption of open-source robotics hardware.
In terms of commercial manufacturers contributing back to the open-source community, Unitree Robotics achieved a major commercial milestone in 2025, shipping more than 5,500 humanoid robots while continuing to open-source its robot SDKs and motion-control environments. Other leading quadruped-robot manufacturers, including DEEP Robotics, have likewise provided open motion-control interfaces to support third-party development. The gradual opening of robot hardware platforms enables the training toolchains, operating systems, and motion-control algorithms described in the previous chapter to be validated and iterated on standardized physical platforms, creating a genuine “software-defined robotics” synergy between hardware and software.
6. Open-Source Embodied AI Tutorials and Talent Development: Bridging Theory and Engineering
Embodied intelligence is a field that places an unusually strong emphasis on systems thinking and real-world engineering. Today, there is a significant mismatch between the available talent pool and the needs of the industry: many existing tutorials focus on algorithmic principles or individual tools, but lack systematic, hands-on projects that span the entire perception–decision-making–control–hardware–debugging pipeline. As a result, many learners emerge as strong in theory but lack the systems-engineering skills required to solve complex real-world problems.
To address this gap, universities, companies, and open-source communities across China have worked closely together to develop a range of high-quality educational projects. According to the “Top 10 EAI Educational Projects of 2025” ranking, the current open-source tutorial ecosystem exhibits three major characteristics:
6.1 Systematic, Full-Stack Learning
The University of Hong Kong’s Embodied-AI-Guide, with more than 12K GitHub stars, has become one of the world’s most influential systematic learning guides for embodied AI.
The Research Direction Starter Guide developed by ScaleLab at Shanghai Jiao Tong University provides newcomers with a clear overview of major research areas and their development trajectories.
The Beijing Innovation Center for Humanoid Robotics has built a comprehensive tutorial system around its HuiSiKaiWu + XR-1 open-source projects, covering the entire workflow from basic development to real-robot deployment.
6.2 Hardware–Software Integration and Platform-Based Hands-On Training
Alibaba DAMO Academy’s Leyun Embodied · SparkEdu leverages a cloud–edge collaborative architecture and comes with an open-source 3D-printed teaching arm, significantly lowering the barriers associated with computing resources and hardware.
OriginBot, an intelligent robotics kit developed by The Constructive Robotics Community (Gouyueju), is built around the ROS ecosystem and has become a popular choice for university education and enterprise prototyping.
Hangzhou Zhigu Future uses ROS as its core technology stack to provide an engineering-oriented training system covering both simulation-based debugging and real-robot deployment.
6.3 Engineering Deployment of Cutting-Edge Algorithms
Unitree RL Training, developed around Unitree’s mainstream robot platforms, has become one of the most widely used reinforcement-learning training tutorials for robotics worldwide.
Hangzhou DEEP Robotics’ Quadruped Motion Control and Embodied Intelligence Development Tutorials provide in-depth coverage of low-level control principles and release the complete source code, with more than one million views across online platforms. Tsinghua AIR’s Embodied Intelligence Reinforcement Learning Bootcamp adopts an immersive theory + simulation + real-robot teaching format.
Zihau Ge’s Robotics and Embodied Intelligence Educational Series has significantly lowered the barrier to entry for the general public, while his LeRobot tutorials have also been incorporated into the official Hugging Face ecosystem.
Through courses, tutorials, hardware platforms, and other educational resources, these projects help learners progress from understanding abstract algorithms to building full-stack robotic systems, creating a positive feedback loop connecting education, talent development, and industry.
Challenges and Outlook
Looking back at 2025, China’s open-source embodied intelligence ecosystem made remarkable progress across models, datasets, evaluation, and foundational infrastructure. Looking ahead, however, the field still faces several significant challenges:
The Generalization Bottleneck: Although current models can perform exceptionally well on specific tasks, their robustness remains limited when confronted with unseen environments, materials, and lighting conditions. There is still a considerable gap between today’s systems and truly general-purpose manipulation that works reliably out of the box.
Gaps in Spatial Perception and Multimodal Understanding: On the one hand, most mainstream VLA models are still derived from 2D vision-language models and lack native understanding of depth, scale, occlusion relationships, and 3D geometric structures. This makes precise localization, grasping, and obstacle avoidance difficult in cluttered and dynamic real-world environments. On the other hand, vision has inherent blind spots under occlusion, in low-light conditions, and during fine-grained contact interactions. Contact-related signals—including force, tactile feedback, and slippage—are critical for tasks such as grasping deformable objects, force-controlled assembly, and dexterous manipulation. Yet the collection of high-quality tactile data and the alignment and fusion of visual and tactile modalities are still at an early stage. Enabling models to both “understand” the three-dimensional structure of physical space and “sense” the forces and textures involved in contact is a critical step toward reliable robotic manipulation.
The Scarcity of High-Quality Data: Although dataset sizes continue to grow, high-quality real-robot data involving complex physical interactions, force feedback, and long-horizon temporal reasoning remains scarce. A structural tension therefore persists between data scale and data quality.
Barriers to Hardware–Software Co-Design: Adapting open-source algorithms to different robot platforms remains costly, while transferring models across embodiments is still difficult. More standardized middleware and interface protocols are urgently needed to reduce these integration costs.
Looking ahead, as reinforcement learning becomes more deeply integrated into embodied intelligence and world models increasingly converge with action models, the capability frontier of embodied AI is likely to continue expanding. Equally important, data-collection methods are becoming increasingly diverse—from real-robot teleoperation and egocentric human demonstrations to the “lifting” of internet video into higher-dimensional training signals and the large-scale generation of synthetic simulation data. These complementary data sources are collectively addressing the long-standing shortage of high-quality embodied data and providing a sustained source of training signal for increasingly capable models. More importantly, these trends are converging toward a deeper transformation: the emergence of truly “embodied-native” foundation models.
As data, algorithms, and hardware continue to evolve in concert, we expect the open-source community to remain a powerful engine of innovation, helping embodied intelligence progress from “usable” to “reliable and effective,” and from task-specific systems to general-purpose capabilities, ultimately bringing artificial general intelligence into the physical world at scale.
References
Ant Group Officially Open-Sources LingBot-Depth, a Next-Generation Spatial Perception Model Based on Masked Depth Modeling. (2026). ModelScope / CSDN
From “Cognition” to “Action”: CogACT Ushers in a New Paradigm for Intelligent Robotic Manipulation. (2025). WeChat Article
AIR Research | X-VLA Open-Sourced, Setting New Records Across Robotic Benchmarks. (2025). Tsinghua AIR
Full-Stack Open Source! Going Beyond π0, a Truly Open-Source Foundation Model for Embodied Intelligence Is Here. (2025). QbitAI
The New Leader in Open-Source Embodied Models? Qianxun Spirit v1.5 Tops RoboChallenge, Bringing the Pi0.5 Era to an End. (2026). Zhihu
ForceRobo Releases DM0, the World’s First Embodied-Native Foundation Model, with the 2.4B-Parameter Version Fully Open-Sourced. (2026). IT Home
The World’s First Autoregressive Video-Action World Model, LingBot-VA, Is Now Open Source. (2026). Zhihu
ShengShu Technology: From Motus to MotuBrain—Ranking First on Both General-Purpose World Action Model Leaderboards. (2026). Tencent News
NVIDIA DreamZero: A Video-Diffusion-Based World Action Model. (2025). Zhihu
Kunlun Tech Officially Open-Sources Matrix-Game: Building a Controllable Interactive World from Images. Yicai. (2025). Yicai
SkyworkAI/Matrix-Game. GitHub. GitHub Repository
Tencent Hunyuan Open-Sources Its Latest World Model, Enabling Real-Time Interactive Generation and Long-Term Spatial Memory. Zhidongxi. (2025). Zhidongxi; Tencent-Hunyuan/HY-WorldPlay. GitHub. GitHub Repository
LingBot-World World Model Officially Open-Sourced. ModelScope Community. (2026). Zhihu; World-Model Open-Source Wave Converges: Ant Group Releases Three Projects in Succession as Google Opens Its Experience. China Daily. (2026). China Daily
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation. (2025). Project Website; arXiv:2506.18088.
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks. Institute for Trustworthy Embodied Intelligence, Fudan University. (2024). VLABench; RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation. (2025).
Open-Source Embodied Intelligence Week: Navigation, Manipulation, Motion Foundation Models and Datasets Released at Scale. Shanghai Artificial Intelligence Laboratory. (2025). Shanghai AI Lab; InternRobotics / EBench. EBench
Shanghai Jiao Tong University Gives Embodied Intelligence a “Standardized Exam”—Could This Become the Robot Equivalent of LMArena? (2026). Zhihu
Based on Tens of Thousands of Real-Robot Evaluations, RoboChallenge Releases Its First Annual Report. (2026). QbitAI
Three Million Sample Pairs, 2.71 TB of Data: Ant Group’s LingBot-Depth-Dataset Open-Sources a Large-Scale Spatial Perception Dataset. (2026). Zhihu; robbyant/LingBot-Depth-Dataset. ModelScope / Hugging Face. Hugging Face Dataset
ModelScope Community & CCF Intelligent Robotics Committee. (2026). 2025 EAI White Paper: Top 10 Educational Projects and Top 10 Datasets. ModelScope
Official Release Information for the RoboMIND Dataset. (2024–2026).
Official Release Information for the AgiBot World Dataset. (2024).
RealMan Open-Sources the World’s First High-Quality Real-Robot Dataset with the Largest Number of Modalities. QbitAI. (2025-11-24). QbitAI; Project Website. RealMan Dataset
EgoLive: A 1,680-Hour Egocentric Dataset of Real-World Tasks. (2026). Zhihu; Dataset. JD Robot Data Market
TASTE-Rob: A Large-Scale Human-Hand Interaction Video Dataset for Generalizable Robotic Manipulation (CVPR 2025). The Chinese University of Hong Kong, Shenzhen. (2025). TASTE-Rob; arXiv
AgiBot Releases the Large-Scale AgiBot Digital World Simulation Framework and Open-Sources a Massive Simulation Dataset. (2025). Zhihu
DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Object Grasping. Peking University (He Wang Research Group). (2023–2025). Project Website; arXiv
Galbot and Tsinghua University Open-Source the OpenWBT Whole-Body Teleoperation System. (2025). QbitAI
Dexbotic Goes Open Source! VLA Performance Improves by 46%, with 100% Success on a Real-Robot Stacking Task. Synced. (2025). Zhihu; Dexmal/dexbotic. GitHub. GitHub Repository
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing. Hong Kong University of Science and Technology and the Open-Source Community. (2026). GitHub Repository; arXiv:2604.05014; Project Website; The “PyTorch Moment” for VLA Has Arrived: HKUST and the Open-Source Community Release StarVLA. Machine Intelligence. (2026-05-09). Zhihu
LeRobot v0.4.0: Major Improvements to Open-Source Robot Learning. (2025). Hugging Face; LeRobot v0.5.0: Scaling Every Dimension. (2026). Hugging Face; huggingface/lerobot. GitHub. GitHub Repository; TheRobotStudio/SO-ARM100. GitHub Repository
RLinf Open-Sourced: The First Large-Scale Reinforcement Learning Framework for Embodied Intelligence with Integrated Rendering, Training, and Inference. Tsinghua University / Zhongguancun Academy / UnknowAI. (2025–2026). Zhihu; RLinf/RLinf. GitHub. GitHub Repository; arXiv:2509.15965.
AgiBot Open-Sources Its Self-Developed Robot Operating System, “Lingqu OS.” IT Home. (2026). Tencent Cloud Developer
Overview of Open-Source Initiatives from Five Major Domestic Embodied-Intelligence Robot Companies (OpenLoong / AimRT). (2025). Zhihu; loongOpen/OpenLoong. GitHub. GitHub Repository
AGIROS Intelligent Robot Operating System Open-Source Community. Institute of Software, Chinese Academy of Sciences. (2025). Institute of Software, CAS; Gitee × AGIROS: Building Domestic Embodied-Intelligence Infrastructure with the Institute of Software, CAS. (2025). Zhihu
dora-rs: Dataflow-Oriented Robotic Architecture. GitHub Repository; Building High-Performance Robotic Applications with dora-rs: From Theory to Practice. (2025). CNBlogs
Giving Robots Instinctive Reactions: Tsinghua Open-Sources a Unified Framework for Parkour and Off-Road Hiking (Project Instinct). BAAI Community. (2026). BAAI Community; A Scalable Perceptive Parkour Framework for Humanoids. arXiv:2601.07718. Project Instinct
Reshaping the R&D Paradigm for Embodied Intelligence: AgiBot Releases Genie Sim 3.0. (2026). Zhihu
The Best Robot Developer Kit Under RMB 1,000 Has Arrived: D-Robotics Launches the RDK X5. (2024). QbitAI
roboto_origin: Fully Open-Source DIY Humanoid Robot. RoboParty. GitHub Repository; A Rising Star in Humanoid Robotics from the HIT Ecosystem: A Fully Open-Source 3 m/s Prototype Developed in Less Than a Year. Ifeng. (2026). Ifeng
DEEP Robotics, the Robot-Dog Maker: What Supports Its 41× Price-to-Sales Ratio? (2026). TMTPost
2025 COSR