en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Embodied AI Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
5
Environment
1
Decision-Making
4
Control
1

100K Video Dialogue Dataset for Real-Time Video Understanding

This dataset contains 100,000 real-time video dialogue sets, each set combines video data with bilingual dialogue transcripts. The dataset covers a diverse range of visual themes, including people, plants, animals, food, and everyday objects. Dialogue topics include factual question answering, contextual questions, follow-up suggestions, and other video-grounded interactions. This dataset can be used for real-time video understanding, vision-language model training, and embodied AI applications.
video dialogue dataset video dialogue data video conversation dataset video dialogue training data multimodal dialogue dataset

100K Egocentric Video Dataset for Procedural Task Understanding

This dataset contains 100,000 egocentric video sets captured from a first-person perspective, covering a diverse range of real-world procedural tasks and interactive activities, including cooking, crafting, sports, and other everyday activities. Annotations include both video-level descriptions and detailed action-level descriptions. This dataset can support the development of video understanding models, egocentric action recognition, vision-language models (VLMs), embodied AI systems, and other AI applications.
egocentric video dataset first-person video dataset procedural video dataset video understanding dataset first-person action dataset video reasoning dataset

10,000-Hour Egocentric Video Dataset for Robotics and AI Manipulation Training

This dataset contains 10,000 hours of egocentric multimodal data collected from diverse real-world environments, including residential, retail, and office scenarios. It covers a wide range of human activities and manipulation tasks, such as meal preparation, cleaning, storage, garment care, merchandising, and object picking. Each sample includes synchronized 4K stereo video, camera calibration parameters, 76-point full-body pose annotations, and fine-grained step-by-step action sequence labels. The dataset is suitable for robot learning, manipulation policy development, and Vision-Language-Action (VLA) models.
embodied ai dataset robotics dataset robot learning dataset robot manipulation dataset vla dataset egocentric video dataset

1,042 Segments 6-camera Egocentric Embodied AI Dataset

This dataset contains 1,042 egocentric video segments (approximately 35 seconds each) collected across 34 locations and 6 real-world environments, including homes, offices, and retail scenarios. Powered by self-developed VSLAM system, achieving millimeter-level positioning, multi-sensor hard-triggered synchronization (≤1ms), and a high frame rate of 60fps for RGB. It includes multi-view videos, calibrations, point clouds, SLAM trajectories, and gesture recognition results in standard formats. Designed for embodied AI, spatial perception, and 3D reconstruction, it offers high precision, diverse scenarios, and out-of-the-box usability, making it ideal for training robust perception-action models.
embodied ai dataset robot learning dataset robotics training data egocentric dataset robot perception dataset multimodal robotics dataset SLAM dataset

116,048 Sets – 3D Hand Pose & Gesture Recognition Dataset

This dataset contains 116,048 sets of 3D handpose data, each set includes hand mask image(RGB, 24-bit), depth image(16-bit), camera intrinsic parameter file(TXT), 3D keypoints file(OBJ), mesh file(OBJ), gesture type file(TXT), keypoints demo image(JPG), and mesh demo image(JPG). The data is collected indoors, with the right hand (no handheld objects), covering both first-person and third-person perspectives, multiple gesture types, finger poses, hand overall rotation poses, individuals and Kinect devices used. This dataset does not include personally identifiable facial information, with hand mask images and depth images aligned. This dataset can be used for tasks such as handpose recognition, hand 3D reconstruction, and hand keypoints detection.
3D hand pose dataset hand keypoints dataset hand gesture recognition RGB-D hand dataset hand mesh dataset AI training hand pose 3D hand reconstruction computer vision hand dataset gesture recognition data

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
23971650-75ca-4a10-bdb0-058ab16331c2