en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Physical AI Data Pyramid: Overview of Data Resources for Nexdata’s Physical AI Data Ecosystem

From:Nexdata Date: 08/14/2026

Recently, Data Pyramid for Embodied Manipulation, a paper published jointly by research teams from Peking University, National Taiwan University, and other universities and institutions, has been gaining attention. In this paper, the Embodied Data Pyramid, which systematically examines the types and functions of training data necessary for Physical AI models along the scalability and robot alignment axes, is proposed.

It is stated in the paper that there is no single optimal data source for Physical AI training. Rather, the effective development of models involves the use of data from several levels of the pyramid. Starting from the top of the pyramid and going downwards, these are Robot Data, UMI Data, Ego Data, Simulation Data, and General Data.

*image from thesis Data Pyramid for Embodied Manipulation

Robot teleoperation data collected at the top level of the pyramid provides the highest degree of robot alignment and can be used for training robotic manipulation skills directly. But the use of physical robotic systems, specialized data collection environments, and complicated data production pipelines makes this type of data expensive and hard to scale.

UMI Data and Ego Data play intermediate roles between robot alignment, scalability, and efficient data collection. They have become crucial data sources for scaling Physical AI model training, while Simulation Data and General Data provide scalable training and a broad knowledge base for environment understanding and further development of model capabilities.

On the basis of the data pyramid, Nexdata is expanding its Physical AI data capabilities at different levels of the pyramid. This gives robotics foundation models data support for human behavior understanding and physical manipulation.

Robot Teleoperation Data

Robot teleoperation data represents the highest level of the Physical AI data pyramid. Collected directly from physical robotic systems, this data includes visual observations, robot states, action trajectories, and feedback from physical interactions during task execution.

10,000 Hours of Dexterous-Hand Teleoperation Data for Physical AI

  • Scale: 10,000 hours, of which 1,000 hours have fine-grained annotations
  • Collection scenarios: simulated workstations, household scenarios, industrial, retail, healthcare scenarios, and more
  • Annotations: task instructions, joint positions, velocity, torque, timestamps, and other parameters

UMI Data

UMI Data represents an important bridge between human demonstrations and robot action learning. Compared with the direct collection of demonstrations on robotic platforms, the use of UMI technology allows the capture of human performance of physical tasks using portable manipulation devices. It makes data collection less dependent on particular robot hardware and less expensive while preserving essential elements of human actions – motion trajectories, action transitions, and object interactions.

These demonstrations allow for effective learning of the pipeline from environment understanding to task execution.

In order to meet the growing demand for UMI Data, Nexdata provides bespoke UMI data collection services, including the full pipeline of task design, environment preparation, data collection, and annotation based on the particular training requirements of the models.

Ego-Centric Data

Ego-Centric Data plays an important role in environment and human behavior understanding. Unlike traditional internet videos, Ego-Centric Data focuses on the actions and interactions of humans in physical environments. This data is captured from a first-person perspective and shows how people perceive their surroundings, navigate environments, interact with objects, and execute tasks. It provides important training signals for spatial reasoning, task understanding, and behavior modeling.

As Physical AI models continue to evolve, Ego-Centric Data is changing too. Traditional Ego-Centric datasets were captured using monocular cameras and focused on the view of the human operator.

But in order to provide strong spatial understanding and task generalization, a single perspective becomes insufficient for representing the complexity of spatial relationships and interactions in the environment. As a result, Ego-Centric Data is transitioning from monocular capture to multi-camera and multi-view configurations. Synchronized multi-camera collection allows the capture of richer environmental context, detailed spatial relationships, and human-object interactions, as well as action transitions across the whole long-horizon task.

Responding to this trend, Nexdata is continuously expanding its Ego-Centric data capabilities across different multi-view configurations.

  • 1,042 Sequences of 6-Camera Multi-View Ego-Centric Data
  • 10,000 Hours of Multi-Scenario Stereo Ego-Centric Data
  • 1,000 Hours of PICO-Based Stereo Physical AI Data
  • 135,000 Hours of Monocular Ego-Centric Data
Scenarios: household, retail, office, and other environments
Data: first-person video, joint parameters, and step-level semantic annotations

Simulation Data

Apart from physically collected data, Simulation Data is another important element of the Physical AI training stack. Training samples can be produced at scale with the help of 3D assets, virtual environments, and physics engines.

Compared to physical data collection, simulation provides controllable environments, high generation efficiency, and repeatable training conditions. It allows for robot environment understanding, task planning, and model generalization. But the Sim-to-Real Gap is still an important problem, making the use of physically collected data indispensable in many cases in order to improve the performance of the model.

In order to support simulation-based training, Nexdata is continuously developing simulation data resources that include 3D models, virtual environments, and other types of data, providing the data foundation for environment understanding, task planning, and robotics model training.

288 Million 3D Models and Scene Assets

  • 3D models: static, interactive, and physics-enhanced models
  • 3D scenes: residential interiors, commercial scenes, and other environments

General Data

General Data includes images, videos, text, and multimodal datasets that allow models to develop visual understanding, language comprehension, and broad world knowledge. Even though this layer does not contain robot actions, it provides an important foundation for robotics foundation models that allows them to understand the environment, objects, instructions, and task semantics.

Nexdata provides multimodal data resources including image, video, speech, and language modalities that support the development of Vision-Language Models (VLMs) and Physical AI models.

100,000 Sets of Real-Time Video Conversation Data

  • Data content: videos accompanied by simulated human-to-human conversations grounded in the video content, with Chinese and English dialogue transcripts
  • Content diversity: plants, animals, food, objects, and other video categories
  • Data features: unlike conventional turn-by-turn Q&A datasets, selected Chinese dialogue transcripts and conversational audio contain interruptions, with<end>used as a marker for interruption points

From the Data Pyramid to Physical AI Data Infrastructure

Data Pyramid for Embodied Manipulation emphasizes an important direction for Physical AI: the next generation of robotics models will not be built on a single data source, but on the use of data at several levels.

Robot teleoperation data provides the highest degree of robot alignment. UMI Data connects human demonstrations with robot action learning. Ego-Centric Data helps models understand human behavior and spatial relationships. Simulation Data provides scalable training in controllable environments, while General Data provides broad world knowledge and multimodal understanding for robotics foundation models.

Together, these various layers form the data foundation for Physical AI.

As this paradigm continues to evolve, Nexdata will continue expanding its Physical AI data capabilities across data collection, data production, and model validation, forming a comprehensive data infrastructure that will support critical stages of robotics model development and enable the next generation of intelligent robots to act autonomously in the physical world.

175f7fe0-36a6-4784-a33a-2c3d59d89172