[{"@type":"PropertyValue","name":"Format","value":"48kHz, 24bit, uncompressed wav, mono channel"},{"@type":"PropertyValue","name":"Recording Environment","value":"Professional recording studio"},{"@type":"PropertyValue","name":"Recording Content","value":"Contains following categories: topic(Extemporaneous speaking on a given topic), multi-level emotion, single-level emotion, paralanguage"},{"@type":"PropertyValue","name":"Personnel","value":"professional voice actors; 1 male, 1 female"},{"@type":"PropertyValue","name":"Annotation Features","value":"Text annotation, emotion annotation, paralinguistic annotation"},{"@type":"PropertyValue","name":"Equipment","value":"Professional recording devices and software"},{"@type":"PropertyValue","name":"Language","value":"Saudi Arabian Arabic"},{"@type":"PropertyValue","name":"Application Scenario","value":"Speech Synthesis"}]
{"id":2244,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_yuyin_default.webp","type1":"165","type1str":null,"type2":"219","type2str":null,"dataname":"2 People - Saudi Arabian Arabic Natural Conversation Average Tone Speech Synthesis Corpus","datazy":[{"title":"Format","content":"48kHz, 24bit, uncompressed wav, mono channel"},{"title":"Recording Environment","content":"Professional recording studio"},{"title":"Recording Content","content":"Contains following categories: topic(Extemporaneous speaking on a given topic), multi-level emotion, single-level emotion, paralanguage"},{"title":"Personnel","content":"professional voice actors; 1 male, 1 female"},{"title":"Annotation Features","content":"Text annotation, emotion annotation, paralinguistic annotation"},{"title":"Equipment","content":"Professional recording devices and software"},{"title":"Language","content":"Saudi Arabian Arabic"},{"title":"Application Scenario","content":"Speech Synthesis"}],"datatag":"Saudi Arabian Arabic,Natural Conversation,Freetalk,Average Tone,TTS,Emotion,Multi-level Emotion,Paralanguage,Interjection","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"F_0001.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260920165125/F_0001.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=j1mQJYC81etSGlzME%2BmlxOwicgk%3D","intro":"إم، إيه والله أنا مثلُك، الدوام هالأسبوع كان ضغط مره، بس جتني فكرة، ليش ما نطلع برا الرياض يوم الخميس؟ نروح الدرعية مثلاً، الجو هناك مره روعة، راح نمشي ونغير جو وناخذ لنا قهوة، والله محتاجين نرتاح بعد هالأسبو .","size":2855188,"progress":100,"type":"mp3"},{"name":"F_0004.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260920165125/F_0004.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=g%2FIEtKPRQMafXhcgLRKV4z0Vp5k%3D","intro":"آه والله يا سلمان، أيام زمان كان لها طعم ثاني، اتذكر طبخ أمنا وشلون كانت تهتم بكل التفاصيل، بس الحين خلينا نحاول نسويها بنفس طريقة، أهم شيء تطلع لذيذة ونفرح فيها.","size":2026273,"progress":100,"type":"mp3"},{"name":"M_0001.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260920165125/M_0001.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=4g3uzQzXc7JauZnPSnBSV9A3cwI%3D","intro":"اممم، إيه والله أنا زيك، الدوام هالأسبوع كان ضغط مره، بس جتني فكرة، ليش ما نطلع برا الرياض يوم الخميس؟ نروح الدرعية مثلاً، الجو هناك حلو ونغير جو شوي، نمشي وناخذ لنا قهوة، والله محتاجين نرتاح بعد هالأسبو .","size":2389063,"progress":100,"type":"mp3"},{"name":"M_0002.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260920165125/M_0002.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=hl78y5BOV8E5NaXoT1Cz7Cf%2Fo8M%3D","intro":"ههه، يا ساتر، انت دايم تفكر بالمواقف، لا تشيل هم، اعرف مكان قريب نوقف فيه، وان شاء الله ما نواجه أي مشكلة، أهم شيء تجي وانت رايق ونستمتع بالطلع .","size":1797592,"progress":100,"type":"mp3"},{"name":"M_0004.wav","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260920165125/M_0004.wav?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=8r%2BLoXZa7E5f5JmX1kTNkFbKP%2B4%3D","intro":"ههه، خلاص اتفقنا، نشوف بعض الساعة سبع الصبح عند محطة البنزين بطريق الملك فهد، أنا جهزت ترمس القهوة، ومرك بالطريق، وإن شاء الله تكون طلعة حلو .","size":1489336,"progress":100,"type":"mp3"}],"officialSummary":"Saudi Arabian Arabic Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.","dataexampl":null,"datakeyword":["Saudi Arabian Arabic","Natural Conversation","Freetalk","Average Tone","TTS","Emotion","Multi-level Emotion","Paralanguage","Interjection"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Voice Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"speechSyn","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"},{\"code\":\"3\",\"language\":\"EN\"}]","productNameEn":"2 People - Saudi Arabian Arabic Natural Conversation Average Tone Speech Synthesis Corpus","BGimg":"brightSpot_audio","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
2 People - Saudi Arabian Arabic Natural Conversation Average Tone Speech Synthesis Corpus
Saudi Arabian Arabic
Natural Conversation
Freetalk
Average Tone
TTS
Emotion
Multi-level Emotion
Paralanguage
Interjection
Saudi Arabian Arabic Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Format
48kHz, 24bit, uncompressed wav, mono channel
Recording Environment
Professional recording studio
Recording Content
Contains following categories: topic(Extemporaneous speaking on a given topic), multi-level emotion, single-level emotion, paralanguage
Personnel
professional voice actors; 1 male, 1 female
Annotation Features
Text annotation, emotion annotation, paralinguistic annotation
Equipment
Professional recording devices and software
Language
Saudi Arabian Arabic
Application Scenario
Speech Synthesis
Sample
Audio
إم، إيه والله أنا مثلُك، الدوام هالأسبوع كان ضغط مره، بس جتني فكرة، ليش ما نطلع برا الرياض يوم الخميس؟ نروح الدرعية مثلاً، الجو هناك مره روعة، راح نمشي ونغير جو وناخذ لنا قهوة، والله محتاجين نرتاح بعد هالأسبو .
Audio
آه والله يا سلمان، أيام زمان كان لها طعم ثاني، اتذكر طبخ أمنا وشلون كانت تهتم بكل التفاصيل، بس الحين خلينا نحاول نسويها بنفس طريقة، أهم شيء تطلع لذيذة ونفرح فيها.
Audio
اممم، إيه والله أنا زيك، الدوام هالأسبوع كان ضغط مره، بس جتني فكرة، ليش ما نطلع برا الرياض يوم الخميس؟ نروح الدرعية مثلاً، الجو هناك حلو ونغير جو شوي، نمشي وناخذ لنا قهوة، والله محتاجين نرتاح بعد هالأسبو .
Audio
ههه، يا ساتر، انت دايم تفكر بالمواقف، لا تشيل هم، اعرف مكان قريب نوقف فيه، وان شاء الله ما نواجه أي مشكلة، أهم شيء تجي وانت رايق ونستمتع بالطلع .
Audio
ههه، خلاص اتفقنا، نشوف بعض الساعة سبع الصبح عند محطة البنزين بطريق الملك فهد، أنا جهزت ترمس القهوة، ومرك بالطريق، وإن شاء الله تكون طلعة حلو .
What languages and voice characteristics are covered by Nexdata’s speech synthesis datasets?
Nexdata offers speech synthesis datasets covering a broad range of languages, dialects, accents, and voice types, supported by extensive global language resources. Our datasets include diverse speakers, speaking styles, emotions, and recording scenarios to support natural and expressive Text-to-Speech (TTS) model development.
Can Nexdata customize speech synthesis datasets for specific languages or requirements?
Yes. If our off-the-shelf TTS datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, annotation, and quality control services. We can customize datasets based on target languages or dialects, speaker profiles, voice characteristics, emotions, speaking styles, recording environments, and data volume.
How does Nexdata ensure the quality and scalability of speech synthesis datasets?
Nexdata applies multi-stage quality control throughout voice data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual TTS projects and specialized datasets for specific voices, accents, emotions, or speech scenarios.