en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Speech Recognition Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Language

All
238
Arabic
7
Burmese
3
Chinese Dialects
3
English
45
French
13
German
11
Hindi
6
Indonesian
8
Italian
11
Japanese
12
Korean
14
Malay
5
Mandarin
3
Others
57
Portuguese
14
Russian
6
Spanish
17
Thai
10
Vietnamese
7

Data Type

All
238
Dialogue
117
Read
122

1013 Hours Brazilian Portuguese Speech Dataset for Speech Recognition

This dataset contains 1,013 hours of Brazilian Portuguese conversational speech, featuring natural conversations and monologues recorded by native speakers from diverse regions. The recordings reflect real-world speaking styles and provide rich linguistic diversity for speech AI development. Each recording includes a transcript, speaker ID, gender, and supporting metadata, enhancing model performance in real and complex tasks. The dataset has undergone quality validation by multiple AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, and PIPL compliant.
brazilian portuguese speech dataset portuguese speech dataset portuguese speech corpus brazilian portuguese asr dataset brazilian portuguese asr dataset

97 Hours Brazilian Portuguese Children Speech Dataset for Speech Recognition

This speech dataset mirrors real-world interactions. Each audio is transcribed and includes text content, speaker ID, gender, age, accent, and other attributes. Our dataset collected from a wide and diverse range of speakers (children aged 12 and under), improves the model's performance on real and complex tasks by geographic location. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Brazilian Portuguese children speech dataset Brazilian Portuguese child speech dataset Brazilian Portuguese children ASR dataset Brazilian Portuguese speech dataset with transcripts

406 Hours European Portuguese Speech Dataset with Spontaneous Dialogues and Transcripts

This dataset contains 406 hours of European Portuguese spontaneous dialogue speech collected through topic-based conversations covering more than 20 domains. Each recording includes accurate transcripts, speaker ID, gender, age, and additional metadata. The dataset was collected from 590 native European Portuguese speakers across diverse regions, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
portuguese audio dataset portuguese speech dataset portuguese voice dataset portuguese speech corpus european portuguese speech dataset portuguese asr dataset

101 Hours Italian Children Speech Dataset for ASR Training

This 101 hours Italian children speech dataset reflects real-world interactions. Each audio sample is transcribed and includes rich metadata such as speaker ID, gender, age, accent, timestamps, and noise-related attributes. Our dataset is collected from a broad and diverse group of speakers (children aged 12 and under) with wide geographical distribution, thus improving the model's performance on real-world, complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Italian children speech dataset Italian child speech corpus Italian kids speech dataset Italian children ASR dataset Italian child voice dataset Italian speech recognition training data

200 Hours - Malay(Malaysia) Spontaneous Dialogue Smartphone speech dataset

Malay(Malaysia) Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics, covering 20+ domains. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(228 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Malay Conversational

162 Hours French Children Speech Dataset for ASR Training

This 162 hours of French children speech dataset mirrors real-world interactions. Each audio sample is transcribed and includes text content, speaker ID, gender, age, accent, and other attributes. Our dataset collected from a wide and diverse range of speakers (children aged 12 and under), improves the model's performance on real and complex tasks by geographic location. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
French children speech dataset French child speech dataset French children ASR dataset French speech dataset

1003 Hours - Hindi Speech Dataset (Spontaneous Conversation)

This dataset contains 1003 hours of Hindi speech audio, mirrors real-world interactions. Each utterance is transcribed with text content, speaker's ID, gender, and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Hindi speech dataset Hindi ASR dataset Hindi TTS dataset Hindi audio dataset Hindi voice dataset

Korean Telephony Speech Dataset – 136 Hours of Spontaneous Calls

This Korean Telephony Speech Dataset contains 136 hours of spontaneous dialogue recorded over phone calls. Covering over 20 real-life domains including customer service, e-commerce, finance, travel, and daily conversations, the dataset features natural two-speaker conversations collected via diverse telephony channels. Each sample is transcribed and annotated with speaker ID, gender, age, and other metadata. Data was collected from 216 native Korean speakers across different regions, enhancing model generalization. Ideal for automatic speech recognition (ASR), speaker diarization, and call center conversational AI systems. All data complies with GDPR, CCPA, and PIPL for responsible and legal AI development.
Korean telephony speech dataset Korean telephone audio telephone conversation Korean call center voice dataset Korean Korean spoken dialogue corpus multilingual telephony dataset Korean voice dataset speech-to-text Korean phone call spontaneous Korean speech data

347 Hours - Indonesian(Indonesia) Spontaneous Dialogue Smartphone speech dataset

Indonesian(Indonesia) Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics, covering 20+ domains. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(412 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Conversational Speech Phone Indonesian
. . .

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
69f09941-7f7c-431a-9707-957a48a0a3c1