en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Home > All Category Datasets > Speech Recognition Datasets > Korean Children Speech Dataset – 393 Hours of Scripted Monologues

Korean Children Speech Dataset – 393 Hours of Scripted Monologues

Korean children speech dataset

Korean child voice dataset

kids speech recognition dataset Korean

Korean ASR training data for children

scripted monologue kids Korean

smartphone voice dataset Korean

Korean kids TTS dataset

Korean educational speech corpus

This 393-hour Korean Children Speech Dataset consists of scripted monologue recordings from young speakers, captured using smartphones. The speech content includes essays, storytelling, and numeric readings. Each audio file is transcribed and annotated with metadata such as speaker ID, gender, and age. Collected from a geographically diverse group of native Korean-speaking children, this dataset is designed to support training of automatic speech recognition (ASR), text-to-speech (TTS), pronunciation evaluation systems, and educational language models. The dataset has been quality-verified by multiple AI enterprises and is fully compliant with GDPR, CCPA, and PIPL privacy regulations.

This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.

Specifications

Specifications

Format

16kHz, 16bit, uncompressed wav, mono channel

Recording environment

quiet indoor environment, without echo

Recording content (read speech)

children's books; human-machine interaction category; smart home command and control category; numbers; general category

Speaker

1,085 Korean children, all children are 6-15 years old

Recording device

Android Smartphone, iPhone

Country

Korea

Language

Korean

Accuracy rate

Sentence Accuracy Rate (SAR) 95%

Sample

Sample

Audio
쁘찌하우스 노부꼬를 예약하고 싶어.
Audio
시간 되면 자주 들어주세요.
Audio
저도 오빠처럼 수영을 잘 하고 싶어요.
Audio
에어컨가 자동으로 켜질수 있게 설정해주세요.
Audio
천이백칠십육만삼천육백십칠원

Recommended Datasets

Recommended Dataset

Infant Laughter Audio Dataset – Baby Laugh Sounds for AI Models

This Infant Laugh Audio Dataset contains 11 minutes of laughter recordings from 20 infants and toddlers aged 0 to 3 years. All audio was captured via smartphone in natural home environments. The dataset supports applications such as baby emotion recognition, laughter detection, and smart parenting AI. Each recording is quality-checked and anonymized. We strictly comply with privacy laws including GDPR, CCPA, and PIPL, ensuring ethical and responsible data usage.

baby laugh dataset infant laughter audio infant emotion dataset baby laugh sound for AI baby voice dataset emotional sound dataset laughter detection audio baby vocal dataset parenting AI dataset infant audio for AI training

Infant Crying Audio Dataset – 52 Hours for AI Baby Cry Detection

Infant Crying smartphone speech dataset, collected by Android smartphone and iPhone, covering infant crying. Our dataset was collected from extensive and diversify speakers(201 people in total, with balanced gender distribution), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

infant crying dataset baby cry detection dataset baby sound dataset crying audio dataset baby monitoring AI parenting AI data smartphone baby cry recording infant audio dataset baby AI training data

50.5 Hours Children Speech Dataset (American English) – Scripted Monologue by Microphone

This dataset contains 50.5 hours of American English speech from children, collected from monologue based on given prompts, covering children’s textbooks, story books, oral language, numbers, letters. Transcribed with text content. Our dataset was collected from extensive and diversify speakers(219 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

child english speech dataset kids speech dataset children speech dataset american english speech dataset american children speech dataset

55 Hours Children Speech Dataset (British English) – Scripted Monologue by Microphone

This dataset contains 55 hours of British English speech from children, collected from monologue based on given scripts, covering educational materials for children, story books, informal language, numbers, alphabet. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(201 British children recorded in hi-fi microphone), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

children speech dataset british english speech dataset british children speech dataset kids speech dataset child english speech dataset

178 Hours - Madarin Chinese(China) Children Scripted Monologue Microphone speech dataset

Madarin Chinese(China) Children Scripted Monologue Microphone speech dataset, collected from monologue based on given scripts, covering educational materials for children, story books, numbers. Transcribed with text content, timestamp and other attributes. Our dataset was collected from extensive and diversify speakers(739 Chinese children recorded in hi-fi microphone), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Chinese children's voice children's data microphone voice acquisition data

Tell Us Your Special Needs

Full Name *

Contact Phone No.*

Company name *

Company Email *

Data Requirements *

By submitting, I agree to the Privacy Protection

Subscribe to our newsletter

Off-the-Shelf Datasets: All Category Datasets; LLM Datasets; Computer Vision Datasets; Speech Recognition Datasets; Speech Synthesis Datasets; OCR Datasets; Pronunciation Dictionary; NLU Datasets

Data Service: 3D Point Cloud Data; Street View Data; OCR Data; Behavior Recognition Data; Identity Recognition Data; Speech Recognition Data; Speech Synthesis Data; Multimodal Data

Industries: Generative AI; Autonomous Vehicles; AR/VR; Conversational AI; Smart Home; Retail; Intelligent Healthcare

Company: About Us; News; Partners; Quality & Security; Event
Links: OPENMPD; DataPlus; Datarade

Platform: Platform
Competition: Competition
Resources: Sponsored Datasets

Sharpen Your AI with Better Data

+1(626)594-5598

[email protected]

nexdata_ai facebook

nexdata_ai twitter

nexdata_ai linkedin

nexdata_ai youtube

Copyright © 2023 NEXDATA TECHNOLOGY INC

Sitemap Terms and Conditions

We use cookies to enhance your browsing experience, serve personalized ads or content, and analyze our traffic. By clicking "Accept All", you consent to our use of cookies.

e23e8bd3-2758-4d27-b719-37702977944b

dfc68574-0b13-4f28-9127-8f029af5c492