234 Hours-Japanese Speech Dataset (Mobile Phone Recordings)

Japanese audio dataset

Japanese ASR training data

Japanese spontaneous dialogue dataset

Japanese speech dataset

This dataset contains 234 hours of Japanese speech audio, collected from monologue based on given scripts, covering 210,000 formal or informal expressions. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(799 Japanese recorded in mixed condition, such as indoor, roadside, restaurant, etc.), geographicly speaking, enhancing model performance in real and complex tasks.Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.

Recommended Dataset

240 Hours - Hindi(India) Speech Dataset (Scripted Monologue)

This dataset collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(401 Indian recorded in quiet and noisy condition), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

hindi phone call speech dataset hindi tts dataset hindi speech corpus hindi audio dataset hindi asr dataset hindi telephony speech dataset hindi dialogue speech dataset hindi conversational speech dataset

227 Hours - Spanish Speech Data by Mobile Phone_R

Spanish Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(352 people in total, from Spain, Mexico and Venezuela), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Spanish Smartphone Reading Scripted Monologue

231.9 Hours - French Scripted Monologue Smartphone speech dataset

French Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(406 speakers, from French, Canada, and Africa etc.), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

French Mobilephone Reading Scripted Monologue

199 Hours - English(the United Kingdom) English Scripted Monologue Smartphone speech dataset

English(the United Kingdom) English Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(346 British people), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

British English pronunciation mobile phone voice data collection voice reading English data

215 Hours - American English Speech Data by Mobile Phone_Reading

English(the United States) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(349 speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

English USA Smartphone Reading Scripted Monologue

127 Hours - Malay(Malaysia) Scripted Monologue Smartphone speech dataset

Malay(Malaysia) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(156 Malaysian), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Malay data mobile phone collected voice data reading voice Malaysian voice

359 Hours - Indonesian(Indonesia) Scripted Monologue Smartphone speech dataset

Indonesian(Indonesia) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering economy, entertainment, news, informal language, numbers, alphabet domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(496 speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Indonesian data mobile phone collected voice data read voice Indonesian voice

203 Hours - Thai(Thailand) Scripted Monologue Smartphone speech dataset

Thai(Thailand) Scripted Monologue Smartphone speech dataset, collected from monologue based on given prompts, covering economy, entertainment, news, oral language, numbers and letters domains. Transcribed with text content. Our dataset was collected from extensive and diversify speakers(498 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Thai Thailand Mobile phone Reading

234 Hours-Japanese Speech Dataset (Mobile Phone Recordings)

Japanese audio dataset Japanese ASR training data Japanese spontaneous dialogue dataset Japanese speech dataset

Current Project Maturity

Japanese audio dataset

Japanese ASR training data

Japanese spontaneous dialogue dataset

Japanese speech dataset