[{"@type":"PropertyValue","name":"Data size","value":"10,000-Hour Egocentric Full-Body Multimodal Dataset"},{"@type":"PropertyValue","name":"Data Distribution","value":"Covers residential, retail & office scenarios (kitchen, bedroom, living room, supermarket, office) with diverse real-life tasks: meal prep, cleaning, storage, garment care, merchandising & picking"},{"@type":"PropertyValue","name":"Data Content","value":"Each sample includes spatiotemporally aligned 4K stereo video, camera calibration params, 76-point full-body pose & step-by-step annotations"},{"@type":"PropertyValue","name":"Capture Solution","value":"Adopts PICO 4 Ultra head-mounted stereo camera + wrist & ankle IMU motion capture solution"},{"@type":"PropertyValue","name":"Data Annotation","value":"Supports dense semantic & action-level annotations; all data passes multi-stage quality control reviews"},{"@type":"PropertyValue","name":"Data Quality","value":"Supports 4096×1536 / 30fps HD video output, tracks 24 torso joints and 52 hand joints, with frame-wise dense annotations and full-process quality control"}]
{"id":2145,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"282","type1str":null,"type2":"283","type2str":null,"dataname":"10,000-Hour Egocentric Video Dataset for Robotics and AI Manipulation Training","datazy":[{"title":"Data size","content":"10,000-Hour Egocentric Full-Body Multimodal Dataset"},{"title":"Data Distribution","content":"Covers residential, retail & office scenarios (kitchen, bedroom, living room, supermarket, office) with diverse real-life tasks: meal prep, cleaning, storage, garment care, merchandising & picking"},{"title":"Data Content","content":"Each sample includes spatiotemporally aligned 4K stereo video, camera calibration params, 76-point full-body pose & step-by-step annotations"},{"title":"Capture Solution","content":"Adopts PICO 4 Ultra head-mounted stereo camera + wrist & ankle IMU motion capture solution"},{"title":"Data Annotation","content":"Supports dense semantic & action-level annotations; all data passes multi-stage quality control reviews"},{"title":"Data Quality","content":"Supports 4096×1536 / 30fps HD video output, tracks 24 torso joints and 52 hand joints, with frame-wise dense annotations and full-process quality control"}],"datatag":" Ego-centric, Embodied AI","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[],"officialSummary":"This dataset contains 10,000 hours of egocentric multimodal data collected from diverse real-world environments, including residential, retail, and office scenarios. It covers a wide range of human activities and manipulation tasks, such as meal preparation, cleaning, storage, garment care, merchandising, and object picking. Each sample includes synchronized 4K stereo video, camera calibration parameters, 76-point full-body pose annotations, and fine-grained step-by-step action sequence labels. The dataset is suitable for robot learning, manipulation policy development, and Vision-Language-Action (VLA) models.","dataexampl":null,"datakeyword":["embodied ai dataset","robotics dataset","robot learning dataset","robot manipulation dataset","vla dataset","egocentric video dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"embodiedAi","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"}]","productNameEn":"10000-Hour Multi-Scenario Egocentric Dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
10,000-Hour Egocentric Video Dataset for Robotics and AI Manipulation Training
embodied ai dataset
robotics dataset
robot learning dataset
robot manipulation dataset
vla dataset
egocentric video dataset
This dataset contains 10,000 hours of egocentric multimodal data collected from diverse real-world environments, including residential, retail, and office scenarios. It covers a wide range of human activities and manipulation tasks, such as meal preparation, cleaning, storage, garment care, merchandising, and object picking. Each sample includes synchronized 4K stereo video, camera calibration parameters, 76-point full-body pose annotations, and fine-grained step-by-step action sequence labels. The dataset is suitable for robot learning, manipulation policy development, and Vision-Language-Action (VLA) models.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Supports dense semantic & action-level annotations; all data passes multi-stage quality control reviews
Data Quality
Supports 4096×1536 / 30fps HD video output, tracks 24 torso joints and 52 hand joints, with frame-wise dense annotations and full-process quality control
What types of Embodied AI applications can Nexdata’s datasets support?
Nexdata’s Embodied AI datasets are designed to support a wide range of robotics and embodied intelligence applications, including robot perception, manipulation, navigation, human-robot interaction, and Vision-Language-Action (VLA) model development. Depending on the dataset, data may include multimodal video, sensor data, robot trajectories, actions, poses, and other annotations for training and evaluation.
Can Nexdata customize Embodied AI datasets based on our specific requirements?
Yes. If our off-the-shelf datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and curation services. We can customize data based on your target robot platform, tasks, environments, sensors, data volume, annotation requirements, and model specifications, helping you build datasets tailored to your specific Embodied AI or VLA project.
Can Nexdata support large-scale Embodied AI data collection projects?
Yes. Nexdata operates an 8,000-square-meter real-world data collection dojo that enables large-scale and diverse data collection for robotics and Embodied AI applications. The facility can support customized environments, task scenarios, robot operations, and multimodal data collection, allowing us to accommodate projects with complex requirements and large data volumes.