178 Hours - Madarin Chinese(China) Children Scripted Monologue Microphone speech dataset

Korean Children Speech Dataset – 393 Hours of Scripted Monologues

This 393-hour Korean Children Speech Dataset consists of scripted monologue recordings from young speakers, captured using smartphones. The speech content includes essays, storytelling, and numeric readings. Each audio file is transcribed and annotated with metadata such as speaker ID, gender, and age. Collected from a geographically diverse group of native Korean-speaking children, this dataset is designed to support training of automatic speech recognition (ASR), text-to-speech (TTS), pronunciation evaluation systems, and educational language models. The dataset has been quality-verified by multiple AI enterprises and is fully compliant with GDPR, CCPA, and PIPL privacy regulations.

Korean children speech dataset Korean child voice dataset kids speech recognition dataset Korean Korean ASR training data for children scripted monologue kids Korean smartphone voice dataset Korean Korean kids TTS dataset Korean educational speech corpus

Infant Laughter Audio Dataset – Baby Laugh Sounds for AI Models

This Infant Laugh Audio Dataset contains 11 minutes of laughter recordings from 20 infants and toddlers aged 0 to 3 years. All audio was captured via smartphone in natural home environments. The dataset supports applications such as baby emotion recognition, laughter detection, and smart parenting AI. Each recording is quality-checked and anonymized. We strictly comply with privacy laws including GDPR, CCPA, and PIPL, ensuring ethical and responsible data usage.

baby laugh dataset infant laughter audio infant emotion dataset baby laugh sound for AI baby voice dataset emotional sound dataset laughter detection audio baby vocal dataset parenting AI dataset infant audio for AI training

Infant Crying Audio Dataset – 52 Hours for AI Baby Cry Detection

Infant Crying smartphone speech dataset, collected by Android smartphone and iPhone, covering infant crying. Our dataset was collected from extensive and diversify speakers(201 people in total, with balanced gender distribution), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

infant crying dataset baby cry detection dataset baby sound dataset crying audio dataset baby monitoring AI parenting AI data smartphone baby cry recording infant audio dataset baby AI training data

50.5 Hours - English(America) Children Scripted Monologue Microphone speech dataset

English(America) Children Scripted Monologue Microphone speech dataset, collected from monologue based on given prompts, covering children’s textbooks, story books, oral language, numbers, letters. Transcribed with text content. Our dataset was collected from extensive and diversify speakers(219 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

US Children's Voice Microphone Acquisition of Voice Data Children's Voice Acquisition data

55 Hours - British Children Speech Data by Microphone

English(the United Kingdom) Children Scripted Monologue Microphone speech dataset, collected from monologue based on given scripts, covering educational materials for children, story books, informal language, numbers, alphabet. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(201 British children recorded in hi-fi microphone), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

English UK British Children Microphone Reading Kid Child Scripted Monologue

178 Hours - Madarin Chinese(China) Children Scripted Monologue Microphone speech dataset

Chinese children's voice children's data microphone voice acquisition data

Chinese children's voice

children's data

microphone voice acquisition data