[{"@type":"PropertyValue","name":"Data Content","value":"1042 segments of multi-scenario ego-centric collected data, each approximately 35 seconds long with complete actions. Each segment includes: 6-camera video, relevant parameters, reconstructed point cloud, SLAM post-processed trajectory, and a video showing gesture keypoint recognition results."},{"@type":"PropertyValue","name":"Data Distribution","value":"Home scenes: 293 segments; Office scenes: 217 segments; Home renovation scenes: 52 segments; Retail scenes: 299 segments; School scenes: 114 segments; Warehouse scenes: 49 segments."},{"@type":"PropertyValue","name":"Data Quality","value":"RGB camera resolution: 1600x1200, frame rate: 60 fps; Monochrome fisheye cameras: 640x480, frame rate: 30 fps. Cameras are hardware-triggered for synchronous exposure."}]
{"id":2236,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"282","type1str":null,"type2":"283","type2str":null,"dataname":"1,042 Segments 6-camera Egocentric Embodied AI Dataset","datazy":[{"title":"Data Content","content":"1042 segments of multi-scenario ego-centric collected data, each approximately 35 seconds long with complete actions. Each segment includes: 6-camera video, relevant parameters, reconstructed point cloud, SLAM post-processed trajectory, and a video showing gesture keypoint recognition results."},{"title":"Data Distribution","content":"Home scenes: 293 segments; Office scenes: 217 segments; Home renovation scenes: 52 segments; Retail scenes: 299 segments; School scenes: 114 segments; Warehouse scenes: 49 segments."},{"title":"Data Quality","content":"RGB camera resolution: 1600x1200, frame rate: 60 fps; Monochrome fisheye cameras: 640x480, frame rate: 30 fps. Cameras are hardware-triggered for synchronous exposure."}],"datatag":"Ego,Embodied AI, 6‑Vision,Physical AI","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[],"officialSummary":"This dataset contains 1,042 egocentric video segments (approximately 35 seconds each) collected across 34 locations and 6 real-world environments, including homes, offices, and retail scenarios. Powered by self-developed VSLAM system, achieving millimeter-level positioning, multi-sensor hard-triggered synchronization (≤1ms), and a high frame rate of 60fps for RGB. It includes multi-view videos, calibrations, point clouds, SLAM trajectories, and gesture recognition results in standard formats. Designed for embodied AI, spatial perception, and 3D reconstruction, it offers high precision, diverse scenarios, and out-of-the-box usability, making it ideal for training robust perception-action models.","dataexampl":null,"datakeyword":["embodied ai dataset","robot learning dataset","robotics training data","egocentric dataset","robot perception dataset","multimodal robotics dataset","SLAM dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"embodiedAi","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"},{\"code\":\"3\",\"language\":\"EN\"}]","productNameEn":"1000 Segments 6-camera Ego Embodied AI Dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
1,042 Segments 6-camera Egocentric Embodied AI Dataset
embodied ai dataset
robot learning dataset
robotics training data
egocentric dataset
robot perception dataset
multimodal robotics dataset
SLAM dataset
This dataset contains 1,042 egocentric video segments (approximately 35 seconds each) collected across 34 locations and 6 real-world environments, including homes, offices, and retail scenarios. Powered by self-developed VSLAM system, achieving millimeter-level positioning, multi-sensor hard-triggered synchronization (≤1ms), and a high frame rate of 60fps for RGB. It includes multi-view videos, calibrations, point clouds, SLAM trajectories, and gesture recognition results in standard formats. Designed for embodied AI, spatial perception, and 3D reconstruction, it offers high precision, diverse scenarios, and out-of-the-box usability, making it ideal for training robust perception-action models.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data Content
1042 segments of multi-scenario ego-centric collected data, each approximately 35 seconds long with complete actions. Each segment includes: 6-camera video, relevant parameters, reconstructed point cloud, SLAM post-processed trajectory, and a video showing gesture keypoint recognition results.
Data Distribution
Home scenes: 293 segments; Office scenes: 217 segments; Home renovation scenes: 52 segments; Retail scenes: 299 segments; School scenes: 114 segments; Warehouse scenes: 49 segments.
Data Quality
RGB camera resolution: 1600x1200, frame rate: 60 fps; Monochrome fisheye cameras: 640x480, frame rate: 30 fps. Cameras are hardware-triggered for synchronous exposure.