[{"@type":"PropertyValue","name":"Data content","value":"100,000 sets of real-time video dialogue data, simulating human-machine conversations based on video content. Each set includes: ①video files (mp4/avi/mov); ②dialogue transcripts in Chinese (json); ③dialogue transcripts in English(json)"},{"@type":"PropertyValue","name":"Data diversity","value":"covers ①video themes(people, plant, animal, food, item, etc); ②dialogue topics(factual QAs, extended suggestions, etc)"},{"@type":"PropertyValue","name":"Data features","value":"incorporates interruptions in selected Chinese dialogue transcripts and audio to better simulate real-world application scenarios"},{"@type":"PropertyValue","name":"Data quality","value":"①video resolution >= 1080p; ②dialogue turns >= 3 rounds per conversation"}]
{"id":1717,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"282","type1str":null,"type2":"283","type2str":null,"dataname":"100,000 Sets - Video Dialogue Dataset","datazy":[{"title":"Data content","content":"100,000 sets of real-time video dialogue data, simulating human-machine conversations based on video content. Each set includes: ①video files (mp4/avi/mov); ②dialogue transcripts in Chinese (json); ③dialogue transcripts in English(json)"},{"title":"Data diversity","content":"covers ①video themes(people, plant, animal, food, item, etc); ②dialogue topics(factual QAs, extended suggestions, etc)"},{"title":"Data features","content":"incorporates interruptions in selected Chinese dialogue transcripts and audio to better simulate real-world application scenarios"},{"title":"Data quality","content":"①video resolution >= 1080p; ②dialogue turns >= 3 rounds per conversation"}],"datatag":"video dialogue,embodied AI","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"7229657.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810103138/7229657.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=1eIxj9G0FuVSemxlqSJOMuWSM4s%3D","intro":"","size":1595140,"progress":100,"type":"jpg"},{"name":"7230113.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810103138/7230113.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=iicXvizuceXLY5y7wpza3JJUoEc%3D","intro":"","size":1818474,"progress":100,"type":"jpg"}],"officialSummary":"100,000 Sets - Video Dialogue Dataset, which includes 100000 videos of multiple themes, with corresponding dialogue text. This dataset can be used for tasks such as video realtime understanding and embodied AI.","dataexampl":null,"datakeyword":["video dialogue","embodied AI"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"embodiedAi","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"100,000 Sets - Video Dialogue Dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
100,000 Sets - Video Dialogue Dataset, which includes 100000 videos of multiple themes, with corresponding dialogue text. This dataset can be used for tasks such as video realtime understanding and embodied AI.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data content
100,000 sets of real-time video dialogue data, simulating human-machine conversations based on video content. Each set includes: ①video files (mp4/avi/mov); ②dialogue transcripts in Chinese (json); ③dialogue transcripts in English(json)