{"id":1119,"datatype":"1","titleimg":"https://res.datatang.com/asset/productNew/APY211101003.png?Expires=2007353700&OSSAccessKeyId=LTAI5tQwXnJZbubgVfVa1ep9&Signature=Y1tOlrWssqZtCevBlrPATEi6kBg%3D","type1":"165","type1str":null,"type2":"166","type2str":null,"dataname":"535 Hours Kazakh Speech Dataset for ASR and Speech Recognition AI Training","datazy":[{"title":"Format","content":"16kHz, 16 bit, wav, mono channel;","desc":"Format"},{"title":"Recording environment","content":"Low background noise;","desc":"Recording environment"},{"title":"Country","content":"China(CHN);","desc":"Country"},{"title":"Language(Region) Code","content":"kk-CN;","desc":"Language(Region) Code"},{"title":"Language","content":"Kazakh;","desc":"Language"},{"title":"Features of annotation","content":"Transcription text, timestamp, speaker ID, gender.","desc":"Features of annotation"},{"title":"Accuracy Rate","content":"Sentence Accuracy Rate (SAR) 95%","desc":"Accuracy Rate"}],"datatag":"Kazakh,Colloquial Video,Conversation","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[],"officialSummary":"This dataset comprises 535 hours of authentic Kazakh speech, featuring recordings of natural conversations and monologues that reflect diverse spoken-language scenarios. Each audio sample is accompanied by an accurate transcript, speaker ID, gender information, and other metadata. Recorded by ethnic Kazakhs from diverse geographic and cultural backgrounds, it offers high accuracy and ease of use, providing a rich resource for speech recognition research and applications while enabling models to perform effectively amidst real-world diversity.. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.","dataexampl":null,"datakeyword":["kazakh speech dataset","kazakh asr dataset","kazakh speech recognition dataset","kazakh voice dataset","low resource speech dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"speechRec","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"535 Hours - Kazakh Spontaneous Speech Data","BGimg":"brightSpot_audio","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
535 Hours Kazakh Speech Dataset for ASR and Speech Recognition AI Training
kazakh speech dataset
kazakh asr dataset
kazakh speech recognition dataset
kazakh voice dataset
low resource speech dataset
This dataset comprises 535 hours of authentic Kazakh speech, featuring recordings of natural conversations and monologues that reflect diverse spoken-language scenarios. Each audio sample is accompanied by an accurate transcript, speaker ID, gender information, and other metadata. Recorded by ethnic Kazakhs from diverse geographic and cultural backgrounds, it offers high accuracy and ease of use, providing a rich resource for speech recognition research and applications while enabling models to perform effectively amidst real-world diversity.. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
This is a paid dataset licensed for commercial use. Ready-made datasets are available for immediate integration into AI projects.