en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

10,000-Hour Egocentric Video Dataset for Robotics and AI Manipulation Training

embodied ai dataset
robotics dataset
robot learning dataset
robot manipulation dataset
vla dataset
egocentric video dataset

This dataset contains 10,000 hours of egocentric multimodal data collected from diverse real-world environments, including residential, retail, and office scenarios. It covers a wide range of human activities and manipulation tasks, such as meal preparation, cleaning, storage, garment care, merchandising, and object picking. Each sample includes synchronized 4K stereo video, camera calibration parameters, 76-point full-body pose annotations, and fine-grained step-by-step action sequence labels. The dataset is suitable for robot learning, manipulation policy development, and Vision-Language-Action (VLA) models.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Data size
10,000-Hour Egocentric Full-Body Multimodal Dataset
Data Distribution
Covers residential, retail & office scenarios (kitchen, bedroom, living room, supermarket, office) with diverse real-life tasks: meal prep, cleaning, storage, garment care, merchandising & picking
Data Content
Each sample includes spatiotemporally aligned 4K stereo video, camera calibration params, 76-point full-body pose & step-by-step annotations
Capture Solution
Adopts PICO 4 Ultra head-mounted stereo camera + wrist & ankle IMU motion capture solution
Data Annotation
Supports dense semantic & action-level annotations; all data passes multi-stage quality control reviews
Data Quality
Supports 4096×1536 / 30fps HD video output, tracks 24 torso joints and 52 hand joints, with frame-wise dense annotations and full-process quality control
Sample Sample
Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

6776df36-ef86-4968-b228-66da5f4e97e3

fd1bee01-226a-40a2-bb57-542e96bfb724