| Dataset Name | Recording Device | Data Size | Specifications |
| 2 Hours- 4 Countries English Speech Synthesis Corpus | Microphone | 2 hours, 4 people | People: 4 people from America, British, Australia, New Zealand Format : 48,000Hz, 24bit, uncompressed wav, mono channel; Recording environment : professional recording studio |
| 20 Hours - France French Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Portugal Language : Portuguese; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - German Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Germany Language : German; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - Italian Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Italy Language : Italian; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - Spain Spanish Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Spain Language : Spanish; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - European Portuguese Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Portugal Language : Portuguese; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - Japanese Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Japan Language : Japanese; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 20 Hours - Korean Reading & Conversational Speech Data by Mobile Phone | Mobile Phone | 20 hours | Format : 16kHz, 16bit, uncompressed wav, mono channel; Recording condition : Low background noise(indoor), without echo; Content category : Reading, Conversation Recording device : Android Smartphone, iPhone; Country : Korea Language : Korean; Features of annotation : Transcription text; Accuracy Rate : Word Accuracy Rate (WAR) is at least 97% |
| 10 Hours - Pashto Conversational Speech Data by Telephone | Telephone | 10 hours | Format : 8kHz 8bit, a-law/u-law pcm, mono channel Content category : Dialogue based on given topics Recording condition : Low background noise (indoor) Recording device : Telephony Country : Afghanistan(AFG) Language(Region) Code : ps-AF Language : Pashto Speaker : 224 people in total, 92% male and 8% female Features of annotation : Transcription text, timestamp, speaker ID, gender Accuracy rate : Word accuracy rate(WAR) 95% Accuracy Rate : Word Accuracy Rate (WAR) is at least 95% |
Note: Please apply for datasets reasonably according to the research field. The maximum number of applications for speech recognition datasets is 4 sets.