[{"@type":"PropertyValue","name":"Data Content","value":"Each participant is associated with two video clips and one TXT document.One video captures the participant reading the content of the TXT document aloud, while maintaining natural lip movements and facial expressions.The other video captures the participant remaining completely silent throughout, with a natural facial expression.The participants are instructed to simulate a video-conference scenario: they look straight at the camera and frame their upper body (from the waist/chest upward)."},{"@type":"PropertyValue","name":"Data Scale","value":"900 persons"},{"@type":"PropertyValue","name":"Gender distribution","value":"Male, Female"},{"@type":"PropertyValue","name":"Ethnic distribution","value":"Asian, White, Black"},{"@type":"PropertyValue","name":"Acquisition environment","value":"Indoor"},{"@type":"PropertyValue","name":"Acquisition devices","value":"Smartphones, webcams (built\u001ein laptop cameras or external USB cameras)"},{"@type":"PropertyValue","name":"Acquisition diversity","value":"Multi-ethnic, multilingual, multi-topic document content, mix of landscape and portrait orientations, multiple backgrounds, multiple age groups"},{"@type":"PropertyValue","name":"Data format","value":"MP4, MOV, and other video formats"},{"@type":"PropertyValue","name":"Video pixel count","value":"The total number of pixels per video is between 777,600 and 8,294,400"},{"@type":"PropertyValue","name":"Accuracy","value":"Label accuracy exceeds 95%; the matching accuracy between the video's read-aloud audio and the document sentences is greater than 95%."}]
{"id":2062,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"220","type2str":null,"dataname":"900-person multi-ethnic facial video collection dataset","datazy":[{"title":"Data Content","content":"Each participant is associated with two video clips and one TXT document.One video captures the participant reading the content of the TXT document aloud, while maintaining natural lip movements and facial expressions.The other video captures the participant remaining completely silent throughout, with a natural facial expression.The participants are instructed to simulate a video-conference scenario: they look straight at the camera and frame their upper body (from the waist/chest upward)."},{"title":"Data Scale","content":"900 persons"},{"title":"Gender distribution","content":"Male, Female"},{"title":"Ethnic distribution","content":"Asian, White, Black"},{"title":"Acquisition environment","content":"Indoor"},{"title":"Acquisition devices","content":"Smartphones, webcams (built\u001ein laptop cameras or external USB cameras)"},{"title":"Acquisition diversity","content":"Multi-ethnic, multilingual, multi-topic document content, mix of landscape and portrait orientations, multiple backgrounds, multiple age groups"},{"title":"Data format","content":"MP4, MOV, and other video formats"},{"title":"Video pixel count","content":"The total number of pixels per video is between 777,600 and 8,294,400"},{"title":"Accuracy","content":"Label accuracy exceeds 95%; the matching accuracy between the video's read-aloud audio and the document sentences is greater than 95%."}],"datatag":"Multi-ethnic,voice, face","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"一组数据.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260819172749/%E4%B8%80%E7%BB%84%E6%95%B0%E6%8D%AE.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=0z3rhL3YLQDQSpbLxq3ydqCeXPk%3D","intro":"","size":9901,"progress":100,"type":"jpg"},{"name":"朗读文档中的内容.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260819172749/%E6%9C%97%E8%AF%BB%E6%96%87%E6%A1%A3%E4%B8%AD%E7%9A%84%E5%86%85%E5%AE%B9.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=QrisgtofwetIJpRXjVbC2YVTg1c%3D","intro":"","size":607776,"progress":100,"type":"jpg"},{"name":"保持沉默.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260819172749/%E4%BF%9D%E6%8C%81%E6%B2%89%E9%BB%98.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=NFt1ADfq9D0CoCjAVwBR69iGNiQ%3D","intro":"","size":587203,"progress":100,"type":"jpg"}],"officialSummary":"900-person multi-ethnic facial video collection dataset,Each participant is associated with two video clips and one TXT document.One video captures the participant reading the content of the TXT document aloud, while maintaining natural lip movements and facial expressions.The other video captures the participant remaining completely silent throughout, with a natural facial expression.The participants are instructed to simulate a video\u001econference scenario: they look straight at the camera and frame their upper body (from the waist/chest upward).Acquisition diversity: Multi-ethnic, multilingual, multi-topic document content, mix of landscape and portrait orientations, multiple backgrounds, multiple age groupsLabel accuracy exceeds 95%; the matching accuracy between the video's read-aloud audio and the document sentences is greater than 95%.The data can be used for application scenarios such as cross-ethnic face recognition, liveness detection, digital humans, and video face swapping.","dataexampl":null,"datakeyword":["Multi-ethnic","voice","face"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Task Type,Modalities","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"computer","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"900-person multi-ethnic facial video collection dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
900-person multi-ethnic facial video collection dataset
Multi-ethnic
voice
face
900-person multi-ethnic facial video collection dataset,Each participant is associated with two video clips and one TXT document.One video captures the participant reading the content of the TXT document aloud, while maintaining natural lip movements and facial expressions.The other video captures the participant remaining completely silent throughout, with a natural facial expression.The participants are instructed to simulate a videoconference scenario: they look straight at the camera and frame their upper body (from the waist/chest upward).Acquisition diversity: Multi-ethnic, multilingual, multi-topic document content, mix of landscape and portrait orientations, multiple backgrounds, multiple age groupsLabel accuracy exceeds 95%; the matching accuracy between the video's read-aloud audio and the document sentences is greater than 95%.The data can be used for application scenarios such as cross-ethnic face recognition, liveness detection, digital humans, and video face swapping.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data Content
Each participant is associated with two video clips and one TXT document.One video captures the participant reading the content of the TXT document aloud, while maintaining natural lip movements and facial expressions.The other video captures the participant remaining completely silent throughout, with a natural facial expression.The participants are instructed to simulate a video-conference scenario: they look straight at the camera and frame their upper body (from the waist/chest upward).
Data Scale
900 persons
Gender distribution
Male, Female
Ethnic distribution
Asian, White, Black
Acquisition environment
Indoor
Acquisition devices
Smartphones, webcams (builtin laptop cameras or external USB cameras)
Acquisition diversity
Multi-ethnic, multilingual, multi-topic document content, mix of landscape and portrait orientations, multiple backgrounds, multiple age groups
Data format
MP4, MOV, and other video formats
Video pixel count
The total number of pixels per video is between 777,600 and 8,294,400
Accuracy
Label accuracy exceeds 95%; the matching accuracy between the video's read-aloud audio and the document sentences is greater than 95%.
What types of computer vision applications can Nexdata’s datasets support?
Nexdata’s computer vision datasets support a wide range of AI applications, including image classification, object detection, image segmentation, facial and human-related recognition, scene understanding, autonomous driving, and other visual perception tasks. Depending on the dataset, data may include images, videos, bounding boxes, polygons, keypoints, segmentation masks, text annotations, and other structured labels.
Can Nexdata customize Computer Vision datasets based on our specific requirements?
Yes. If our off-the-shelf Computer Vision datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and curation services. We can customize data based on your target objects, environments, scenarios, camera specifications, geographic locations, data volume, annotation formats, and quality standards to support specific model training and evaluation needs.
How does Nexdata ensure the quality and scalability of its Computer Vision datasets?
Nexdata applies multi-stage quality control throughout data collection, annotation, validation, and delivery. Depending on project requirements, we can implement customized annotation guidelines, multi-level reviews, consistency checks, and quality sampling to ensure dataset accuracy and consistency. Our data collection and processing capabilities can also be scaled to support large-volume Computer Vision projects.