[{"@type":"PropertyValue","name":"Format","value":"16kHz, 16 bit, wav, mono channel;"},{"@type":"PropertyValue","name":"Content category","value":"Recorders in free conversation without a set topic;"},{"@type":"PropertyValue","name":"Recording condition","value":"Low background noise (indoor);"},{"@type":"PropertyValue","name":"Recording device","value":"Android smartphone, iPhone;"},{"@type":"PropertyValue","name":"Speaker","value":"268 native speakers in total, 41% male and 59% female;"},{"@type":"PropertyValue","name":"Country","value":"Kingdom of Saudi Arabia;"},{"@type":"PropertyValue","name":"Language","value":"Arabic;"},{"@type":"PropertyValue","name":"Features of annotation","value":"Transcription text, timestamp, speaker ID, gender;"},{"@type":"PropertyValue","name":"Accuracy Rate","value":"Word Accuracy Rate (WAR) 95%;"}]
{"id":1627,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_yuyin_default.webp","type1":"165","type1str":null,"type2":"170","type2str":null,"dataname":"268 Hours Arabic Full-Duplex Speech Dataset with Multi-Channel Conversations","datazy":[{"title":"Format","content":"16kHz, 16 bit, wav, mono channel;"},{"title":"Content category","content":"Recorders in free conversation without a set topic;"},{"title":"Recording condition","content":"Low background noise (indoor);"},{"title":"Recording device","content":"Android smartphone, iPhone;"},{"title":"Speaker","content":"268 native speakers in total, 41% male and 59% female;"},{"title":"Country","content":"Kingdom of Saudi Arabia;"},{"title":"Language","content":"Arabic;"},{"title":"Features of annotation","content":"Transcription text, timestamp, speaker ID, gender;"},{"title":"Accuracy Rate","content":"Word Accuracy Rate (WAR) 95%;"}],"datatag":" Arabic,Multi-stream, Dialogue,full duplex","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"0001_001_A-1.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20250702171153/0001_001_A-1.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=imdGiRsurWjnHPKBsgEHNpsJL6I%3D","intro":"هل في طريقة أزيد فيها مستوى التأمين على حسابي؟","size":139500,"progress":100,"type":"mp3"},{"name":"0001_001_A-2.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20250702171153/0001_001_A-2.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=fGpuyWAF%2BLtNZZbzQe1uUhRMcJI%3D","intro":"وإذا كنت أحتاج مستند يوضح تفاصيل التأمين لحساباتي، كيف أقدر أحصله؟","size":227916,"progress":100,"type":"mp3"},{"name":"0001_001_A-3.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20250702171153/0001_001_A-3.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=VW6uHYFZCle1r0SbW5TeuHjTYO4%3D","intro":"طيب وش الإجراءات اللي تتم في حال صار أي خلل في البنك، كيف أقدر استرجع فلوسي؟","size":266764,"progress":100,"type":"mp3"},{"name":"0001_001_A-4.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20250702171153/0001_001_A-4.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=epYMV%2FS61ySVQ4qMDoVy6g54634%3D","intro":"يعني ما يحتاج أقدم طلب وأتابع الموضوع بنفسي؟","size":138796,"progress":100,"type":"mp3"},{"name":"0001_001_A-5.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20250702171153/0001_001_A-5.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=SyabuQS0WpcIkm8n%2FEEcLbROitg%3D","intro":"تمام، بخصوص الحسابات اللي مسجل فيها أكثر من مستفيد، كيف يتم التعامل معها في التأمين؟","size":238604,"progress":100,"type":"mp3"}],"officialSummary":"This dataset features full-duplex, multi-channel conversations recorded from 268 native Saudi Arabic speakers in authentic customer service scenarios. The dataset includes high-quality audio with verbatim transcriptions and comprehensive metadata, such as speaker ID, gender, age, and other demographic attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.","dataexampl":null,"datakeyword":["full-duplex speech dataset","multi-channel audio dataset","Saudi Arabic speech dataset","Arabic customer service speech dataset","Arabic speech dataset","Saudi Arabic speech dataset","Arabic call center dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"speechRec","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"268 Hours - Arabic(Saudi) Full-Duplex Spontaneous Dialogue Smartphone speech dataset-Customer Service","BGimg":"brightSpot_audio","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
268 Hours Arabic Full-Duplex Speech Dataset with Multi-Channel Conversations
full-duplex speech dataset
multi-channel audio dataset
Saudi Arabic speech dataset
Arabic customer service speech dataset
Arabic speech dataset
Saudi Arabic speech dataset
Arabic call center dataset
This dataset features full-duplex, multi-channel conversations recorded from 268 native Saudi Arabic speakers in authentic customer service scenarios. The dataset includes high-quality audio with verbatim transcriptions and comprehensive metadata, such as speaker ID, gender, age, and other demographic attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Format
16kHz, 16 bit, wav, mono channel;
Content category
Recorders in free conversation without a set topic;
Recording condition
Low background noise (indoor);
Recording device
Android smartphone, iPhone;
Speaker
268 native speakers in total, 41% male and 59% female;
What languages and scenarios are covered by Nexdata’s speech recognition datasets?
Nexdata offers speech recognition datasets covering a broad range of languages, dialects, and accents, backed by extensive global language resources. Our datasets include diverse speakers, acoustic environments, and real-world speech scenarios, supporting multilingual ASR, voice assistants, conversational AI, speech-to-text, and other speech-enabled applications.
Can Nexdata customize speech recognition datasets for specific languages or requirements?
Yes. If our off-the-shelf datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, and annotation services. We can customize datasets based on target languages or dialects, speaker profiles, recording environments, speech scenarios, data volume, and annotation specifications to meet specific ASR development needs.
How does Nexdata ensure the quality and scalability of speech recognition datasets?
Nexdata applies multi-stage quality control throughout speech data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual projects and specialized datasets for specific languages, dialects, accents, and speech scenarios.