600 Hours Greek Speech Dataset – Real world Casual Conversation & Monologue for ASR

INTERSPEECH 2025 MLC-SLM Challenge Dataset

The INTERSPEECH 2025 MLC-SLM Challenge Dataset, curated by Nexdata, is derived from fifteen proprietary conversational speech corpora. Distinguished by exceptional annotation accuracy and operational reliability, this dataset is engineered to address critical challenges in multilingual automatic speech recognition (ASR) and long-context comprehension. It meticulously replicates real-world complexities including spontaneous interruptions and speaker overlaps across 11 languages (1500 hours total duration), thereby providing robust training resources for developing world-ready ASR systems. All data collection and processing strictly comply with international privacy regulations including GDPR, CCPA and PIPL, with rigorous protocols ensuring participant anonymity and ethical data usage throughout the lifecycle.

workshop audio dataset mlc-slm dataset ASR speech recognition data

3000 Hours - Mandarin Full-Duplex Spontaneous Dialogue Speech Dataset

Mandarin Full-Duplex Spontaneous Dialogue Speech Dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Full-Duplex Dialogue Mandarin

600 Hours Norwegian Speech Dataset – Real-world Casual Conversation & Monologue for ASR

The 600 Hours Norwegian Real-World Speech Dataset includes both casual conversations and monologues, covering domains such as self-media, live shows, and other generic domains, mirroring real-world interactions. Transcribed with text content, speaker's ID, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

norwegian speech dataset norwegian ASR training data norwegian conversation corpus norwegian monologue speech norwegian speech recognition dataset speech-to-text norwegian data norwegian voice dataset multilingual speech data norwegian transcription dataset

1300 Hours Gujatati(India) Speech Dataset (Scripted Dialogue)

This dataset contains 1,300 hours of Gujarati speech, covers several domains, mirrors real-world interactions. Transcribed with text content, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

gujarati audio dataset gujarati asr dataset gujarati speech dataset gujarati tts dataset

Spanish(Mexico) Real-world Casual Conversation and Monologue speech dataset

Spanish(Mexico) Real-world Casual Conversation and Monologue speech dataset, covers self-media, conversation, variety show and other generic domains, mirrors real-world interactions. Transcribed with text content, speaker's ID, gender, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Mexico Spanish Casual Conversation ASR

460 Hour Swedish Speech Dataset - Casual Conversations & Monologues for ASR & TTS

This dataset contains 460 hours of Swedish speech, mirrors real-world interactions. Transcribed with text content, speaker's ID, gender, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Swedish speech dataset Swedish conversation corpus Real-world Swedish audio Swedish ASR training data TTS training dataset Swedish Swedish NLP corpus

1,503 Hours – UAE Arabic Speech Dataset for TTS & ASR

This dataset mirrors real-world interactions. Transcribed with text content, speaker's ID, gender, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

UAE Arabic Speech Dataset Arabic speech dataset Arabic Conversational Dataset Arabic Speech Corpus Arabic monologue speech data

1500 Hours - French(Canada) Real-world Casual Conversation and Monologue speech dataset

French(Canada) Real-world Casual Conversation and Monologue speech dataset, covers self-media, conversation, variety show and other generic domains, mirrors real-world interactions. Transcribed with text content, speaker's ID, gender, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Canada French Casual Conversation ASR

600 Hours Greek Speech Dataset – Real world Casual Conversation & Monologue for ASR

greek speech dataset greek ASR training data greek conversation corpus greek monologue speech greek speech recognition dataset speech-to-text greek data greek voice dataset greek transcription dataset

greek speech dataset

greek ASR training data

greek conversation corpus

greek monologue speech

greek speech recognition dataset

speech-to-text greek data

greek voice dataset

greek transcription dataset