[{"@type":"PropertyValue","name":"Data content","value":"101,759 video QA sets, each contains one video and its corresponding Q&A annotation file."},{"@type":"PropertyValue","name":"Data distribution","value":"①10 video genres: office, education, travel, cooking, product, news, fashion, sports, health, entertainment; ②multiple question types: action, action-change, location, location-change, scene, scene-change, attribute, attribute-change, etc."},{"@type":"PropertyValue","name":"Annotation","value":"Q&A turns grounded in the video. Each round includes a question, an answer, a flag indicating whether audio is required, and start/end time-stamps."},{"@type":"PropertyValue","name":"Data quality","value":"①video: >60s, >=1080p, with audio; ②annotation: >=5 turns per video, >=10 characters per turn for both question and answer; ③accuracy: >=95% of all Q&A turns are correct."}]
{"id":1827,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"226","type1str":null,"type2":"254","type2str":null,"dataname":"101759 Sets - Video QA Dataset","datazy":[{"title":"Data content","content":"101,759 video QA sets, each contains one video and its corresponding Q&A annotation file."},{"title":"Data distribution","content":"①10 video genres: office, education, travel, cooking, product, news, fashion, sports, health, entertainment; ②multiple question types: action, action-change, location, location-change, scene, scene-change, attribute, attribute-change, etc."},{"title":"Annotation","content":"Q&A turns grounded in the video. Each round includes a question, an answer, a flag indicating whether audio is required, and start/end time-stamps."},{"title":"Data quality","content":"①video: >60s, >=1080p, with audio; ②annotation: >=5 turns per video, >=10 characters per turn for both question and answer; ③accuracy: >=95% of all Q&A turns are correct."}],"datatag":"VQA,multi-modal","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"0000001.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810153717/0000001.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=Pq0zVDxHCqV9V%2FlTrRDnVryKzmA%3D","intro":"","size":184820,"progress":100,"type":"jpg"},{"name":"0050936.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810153717/0050936.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=98ievKBbPwZuFvh%2Ff6%2BzRJ617IQ%3D","intro":"","size":180789,"progress":100,"type":"jpg"},{"name":"0081739.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810153717/0081739.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=G1d4KPEwWtw4Nq%2BGqx1q0EKMJIQ%3D","intro":"","size":214122,"progress":100,"type":"jpg"}],"officialSummary":"101759 Sets - Video QA Dataset, which include 10 genres of video and its corresponding QA annotation. The 10 genres are office/education/travel/cooking/product/news/fashion/sports/health/entertainment. Videos are >60s, >=1080p, with audio. QA annotations have multiple question types, like action/action-change/location/location-change/scene/scene-change/attribute/attribute-change, etc. QA annotations are >=5 turns per video, >=10 characters per turn for both question and answer. This dataset can be used for multi-modal understanding tasks, like VQA, etc.","dataexampl":null,"datakeyword":["VQA","multi-modal"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"llm","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"101759 Sets - Video QA Dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
101759 Sets - Video QA Dataset, which include 10 genres of video and its corresponding QA annotation. The 10 genres are office/education/travel/cooking/product/news/fashion/sports/health/entertainment. Videos are >60s, >=1080p, with audio. QA annotations have multiple question types, like action/action-change/location/location-change/scene/scene-change/attribute/attribute-change, etc. QA annotations are >=5 turns per video, >=10 characters per turn for both question and answer. This dataset can be used for multi-modal understanding tasks, like VQA, etc.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data content
101,759 video QA sets, each contains one video and its corresponding Q&A annotation file.
Q&A turns grounded in the video. Each round includes a question, an answer, a flag indicating whether audio is required, and start/end time-stamps.
Data quality
①video: >60s, >=1080p, with audio; ②annotation: >=5 turns per video, >=10 characters per turn for both question and answer; ③accuracy: >=95% of all Q&A turns are correct.