[{"@type":"PropertyValue","name":"Format","value":"16kHz, 16bit, uncompressed wav, mono channel"},{"@type":"PropertyValue","name":"Recording environment","value":"quiet indoor environment, without echo"},{"@type":"PropertyValue","name":"Recording content (read speech)","value":"children's books and textbooks"},{"@type":"PropertyValue","name":"Speaker","value":"286 American children, 53% of which are female, all children are 5-12 years old"},{"@type":"PropertyValue","name":"Recording device","value":"Android smartphone, iPhone"},{"@type":"PropertyValue","name":"Country","value":"The United States of America(USA)"},{"@type":"PropertyValue","name":"Language","value":"English"},{"@type":"PropertyValue","name":"Language(Region) Code","value":"en-US"},{"@type":"PropertyValue","name":"Accuracy rate","value":"Sentence accuracy rate(SAR) 95%"}]
{"id":1197,"datatype":"1","titleimg":"https://res.datatang.com/asset/productNew/APY221015001.jpg?Expires=2007353715&OSSAccessKeyId=LTAI5tQwXnJZbubgVfVa1ep9&Signature=5SvsGM7rIIKpHgJjZxXddkS003E%3D","type1":"165","type1str":null,"type2":"167","type2str":null,"dataname":"299 Hours – US English Children Speech Dataset for ASR & TTS","datazy":[{"title":"Format","content":"16kHz, 16bit, uncompressed wav, mono channel","desc":"Format"},{"title":"Recording environment","content":"quiet indoor environment, without echo","desc":"Recording environment"},{"title":"Recording content (read speech)","content":"children's books and textbooks","desc":"Recording content (read speech)"},{"title":"Speaker","content":"286 American children, 53% of which are female, all children are 5-12 years old","desc":"Speaker"},{"title":"Recording device","content":"Android smartphone, iPhone","desc":"Recording device"},{"title":"Country","content":"The United States of America(USA)","desc":"Country"},{"title":"Language","content":"English","desc":"Language"},{"title":"Language(Region) Code","content":"en-US","desc":"Language(Region) Code"},{"title":"Accuracy rate","content":"Sentence accuracy rate(SAR) 95%","desc":"Accuracy rate"}],"datatag":"English,Children,American,American English","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[{"name":"/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00102S0046.wav","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00102S0046.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=Aoo0jntJrlQZbfJghrEbVFDmJsY%3D","intro":"he wore baggy trousers and a long shirt, his face was almost completely hidden by his head cloth. he did not speak or look at them.","size":0,"progress":100,"type":"mp3"},{"name":"/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00114S0028.wav","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00114S0028.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=h2b6S%2BNplgFDROVU%2F1Qoqv%2B8U2Y%3D","intro":"for his or her intelligence. that goes for the animals as well as the people. everything that happens to them is explained to us.","size":0,"progress":100,"type":"mp3"},{"name":"/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00154S0042.wav","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00154S0042.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=H%2Fe6D%2FZVZ7rCzByh%2FrPASfq%2FkpA%3D","intro":"arabia where the greatest horses in the world were bred!","size":0,"progress":100,"type":"mp3"},{"name":"/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00010S0002.wav","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00010S0002.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=4tL8dWJZl2v5%2BDS%2BB%2FsE5NKsKKY%3D","intro":"chapter three of terror at the zoo by peg kehret.","size":0,"progress":100,"type":"mp3"},{"name":"/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00163S0008.wav","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY221015001_demo1712570408477/G00163S0008.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=aZOEp9bXqxzI%2FSYlxa2nLLlKAII%3D","intro":"jack pushed his glasses into place. who was going to believe any","size":0,"progress":100,"type":"mp3"}],"officialSummary":"This dataset includes 299 hours of US English children’s speech, recorded as scripted monologues, collected from monologue based on given scripts, covering essay stories. The data covers a variety of categories, including children's books and textbooks, and is rich in content that aligns with children's language habits.Transcribed with text content and other attributes. Our dataset is collected from a wide and diverse range of speakers geographically, which supports tasks like speech recognition, TTS, and child voice modeling. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.","dataexampl":null,"datakeyword":["children speech dataset","kids voice dataset","child speech corpus","US English children dataset","children TTS dataset","child speech recognition data","American English children voice dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"speechRec","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"299 Hours - American Children Speech Data By Mobile Phone","BGimg":"brightSpot_audio","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
299 Hours – US English Children Speech Dataset for ASR & TTS
children speech dataset
kids voice dataset
child speech corpus
US English children dataset
children TTS dataset
child speech recognition data
American English children voice dataset
This dataset includes 299 hours of US English children’s speech, recorded as scripted monologues, collected from monologue based on given scripts, covering essay stories. The data covers a variety of categories, including children's books and textbooks, and is rich in content that aligns with children's language habits.Transcribed with text content and other attributes. Our dataset is collected from a wide and diverse range of speakers geographically, which supports tasks like speech recognition, TTS, and child voice modeling. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
This is a paid dataset licensed for commercial use. Ready-made datasets are available for immediate integration into AI projects.
Specifications
Format
16kHz, 16bit, uncompressed wav, mono channel
Recording environment
quiet indoor environment, without echo
Recording content (read speech)
children's books and textbooks
Speaker
286 American children, 53% of which are female, all children are 5-12 years old
Recording device
Android smartphone, iPhone
Country
The United States of America(USA)
Language
English
Language(Region) Code
en-US
Accuracy rate
Sentence accuracy rate(SAR) 95%
Sample
Audio
he wore baggy trousers and a long shirt, his face was almost completely hidden by his head cloth. he did not speak or look at them.
Audio
for his or her intelligence. that goes for the animals as well as the people. everything that happens to them is explained to us.
Audio
arabia where the greatest horses in the world were bred!
Audio
chapter three of terror at the zoo by peg kehret.
Audio
jack pushed his glasses into place. who was going to believe any
What languages and scenarios are covered by Nexdata’s speech recognition datasets?
Nexdata offers speech recognition datasets covering a broad range of languages, dialects, and accents, backed by extensive global language resources. Our datasets include diverse speakers, acoustic environments, and real-world speech scenarios, supporting multilingual ASR, voice assistants, conversational AI, speech-to-text, and other speech-enabled applications.
Can Nexdata customize speech recognition datasets for specific languages or requirements?
Yes. If our off-the-shelf datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, and annotation services. We can customize datasets based on target languages or dialects, speaker profiles, recording environments, speech scenarios, data volume, and annotation specifications to meet specific ASR development needs.
How does Nexdata ensure the quality and scalability of speech recognition datasets?
Nexdata applies multi-stage quality control throughout speech data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual projects and specialized datasets for specific languages, dialects, accents, and speech scenarios.