From:Nexdata Date: 08/14/2026
Recently, Data Pyramid for Embodied Manipulation, a paper published jointly by research teams from Peking University, National Taiwan University, and other universities and institutions, has been gaining attention. In this paper, the Embodied Data Pyramid, which systematically examines the types and functions of training data necessary for Physical AI models along the scalability and robot alignment axes, is proposed.
It is stated in the paper that there is no single optimal data source for Physical AI training. Rather, the effective development of models involves the use of data from several levels of the pyramid. Starting from the top of the pyramid and going downwards, these are Robot Data, UMI Data, Ego Data, Simulation Data, and General Data.
*image from thesis Data Pyramid for Embodied Manipulation
Robot teleoperation data collected at the top level of the pyramid provides the highest degree of robot alignment and can be used for training robotic manipulation skills directly. But the use of physical robotic systems, specialized data collection environments, and complicated data production pipelines makes this type of data expensive and hard to scale.
UMI Data and Ego Data play intermediate roles between robot alignment, scalability, and efficient data collection. They have become crucial data sources for scaling Physical AI model training, while Simulation Data and General Data provide scalable training and a broad knowledge base for environment understanding and further development of model capabilities.
On the basis of the data pyramid, Nexdata is expanding its Physical AI data capabilities at different levels of the pyramid. This gives robotics foundation models data support for human behavior understanding and physical manipulation.
Robot teleoperation data represents the highest level of the Physical AI data pyramid. Collected directly from physical robotic systems, this data includes visual observations, robot states, action trajectories, and feedback from physical interactions during task execution.
10,000 Hours of Dexterous-Hand Teleoperation Data for Physical AI
UMI Data
UMI Data represents an important bridge between human demonstrations and robot action learning. Compared with the direct collection of demonstrations on robotic platforms, the use of UMI technology allows the capture of human performance of physical tasks using portable manipulation devices. It makes data collection less dependent on particular robot hardware and less expensive while preserving essential elements of human actions – motion trajectories, action transitions, and object interactions.
These demonstrations allow for effective learning of the pipeline from environment understanding to task execution.
In order to meet the growing demand for UMI Data, Nexdata provides bespoke UMI data collection services, including the full pipeline of task design, environment preparation, data collection, and annotation based on the particular training requirements of the models.
Ego-Centric Data plays an important role in environment and human behavior understanding. Unlike traditional internet videos, Ego-Centric Data focuses on the actions and interactions of humans in physical environments. This data is captured from a first-person perspective and shows how people perceive their surroundings, navigate environments, interact with objects, and execute tasks. It provides important training signals for spatial reasoning, task understanding, and behavior modeling.
As Physical AI models continue to evolve, Ego-Centric Data is changing too. Traditional Ego-Centric datasets were captured using monocular cameras and focused on the view of the human operator.
But in order to provide strong spatial understanding and task generalization, a single perspective becomes insufficient for representing the complexity of spatial relationships and interactions in the environment. As a result, Ego-Centric Data is transitioning from monocular capture to multi-camera and multi-view configurations. Synchronized multi-camera collection allows the capture of richer environmental context, detailed spatial relationships, and human-object interactions, as well as action transitions across the whole long-horizon task.
Responding to this trend, Nexdata is continuously expanding its Ego-Centric data capabilities across different multi-view configurations.
Simulation Data
Apart from physically collected data, Simulation Data is another important element of the Physical AI training stack. Training samples can be produced at scale with the help of 3D assets, virtual environments, and physics engines.
Compared to physical data collection, simulation provides controllable environments, high generation efficiency, and repeatable training conditions. It allows for robot environment understanding, task planning, and model generalization. But the Sim-to-Real Gap is still an important problem, making the use of physically collected data indispensable in many cases in order to improve the performance of the model.
In order to support simulation-based training, Nexdata is continuously developing simulation data resources that include 3D models, virtual environments, and other types of data, providing the data foundation for environment understanding, task planning, and robotics model training.
288 Million 3D Models and Scene Assets
General Data includes images, videos, text, and multimodal datasets that allow models to develop visual understanding, language comprehension, and broad world knowledge. Even though this layer does not contain robot actions, it provides an important foundation for robotics foundation models that allows them to understand the environment, objects, instructions, and task semantics.
Nexdata provides multimodal data resources including image, video, speech, and language modalities that support the development of Vision-Language Models (VLMs) and Physical AI models.
100,000 Sets of Real-Time Video Conversation Data
Data Pyramid for Embodied Manipulation emphasizes an important direction for Physical AI: the next generation of robotics models will not be built on a single data source, but on the use of data at several levels.
Robot teleoperation data provides the highest degree of robot alignment. UMI Data connects human demonstrations with robot action learning. Ego-Centric Data helps models understand human behavior and spatial relationships. Simulation Data provides scalable training in controllable environments, while General Data provides broad world knowledge and multimodal understanding for robotics foundation models.
Together, these various layers form the data foundation for Physical AI.
As this paradigm continues to evolve, Nexdata will continue expanding its Physical AI data capabilities across data collection, data production, and model validation, forming a comprehensive data infrastructure that will support critical stages of robotics model development and enable the next generation of intelligent robots to act autonomously in the physical world.