[{"@type":"PropertyValue","name":"Format","value":"48,000Hz, 24bit, uncompressed wav, mono channel;"},{"@type":"PropertyValue","name":"Recording environment","value":"professional recording studio;"},{"@type":"PropertyValue","name":"Recording content","value":"general narrative sentences, interrogative sentences, etc;"},{"@type":"PropertyValue","name":"Speaker","value":"male, 20-30 years old, young and positive voice;"},{"@type":"PropertyValue","name":"Device","value":"microphone;"},{"@type":"PropertyValue","name":"Language","value":"American English;"},{"@type":"PropertyValue","name":"Annotation","value":"word and phoneme transcription, four-level prosodic boundary annotation;"},{"@type":"PropertyValue","name":"Application scenarios","value":"speech synthesis."}]
{"id":1159,"datatype":"1","titleimg":"https://res.datatang.com/asset/productNew/APY220430001.png?Expires=2007353707&OSSAccessKeyId=LTAI5tQwXnJZbubgVfVa1ep9&Signature=Xy0LTK6smT2ZIR6bTMvkcM%2Bj/c0%3D","type1":"165","type1str":null,"type2":"165","type2str":null,"dataname":"20 Hours - American English Speech Synthesis Corpus-Male","datazy":[{"title":"Format","value":"48,000Hz, 24bit, uncompressed wav, mono channel;"},{"title":"Recording environment","value":"professional recording studio;"},{"title":"Recording content","value":"general narrative sentences, interrogative sentences, etc;"},{"title":"Speaker","value":"male, 20-30 years old, young and positive voice;"},{"title":"Device","value":"microphone;"},{"title":"Language","value":"American English;"},{"title":"Annotation","value":"word and phoneme transcription, four-level prosodic boundary annotation;"},{"title":"Application scenarios","value":"speech synthesis."}],"datatag":"English,Tts,American English,Male","technologydoc":null,"downurl":null,"datainfo":"","standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[["mp3","https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100003.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=EsLOpjsxcnlIoj4qqwkJbQ1TajY%3D","/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100003.wav","Look- at- the way they hear operas- and- see oil paintings%.L UH1 K3 / AE1 T / DH AX0 / W EY1 / DH EY1 / HH IY1 R / AA1 . P R AX0 Z / AX0 N D / S IY1 / OY1 L / P EY1 N . T IH0 NG Z"],["mp3","https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100009.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=1zsMh1kphe0l7IMctfv5kpDKjso%3D","/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100009.wav","Was there some discussion- about- whether I should- speak%?W AX1 Z / DH EH1 R / S AH1 M / D IH0 . S K AH1 . SH AX0 N3 / AX0 . B AW1 T / W EH1 . DH ER0 / AY1 / SH UH1 D / S P IY1 K"],["mp3","https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100005.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=5Qdf67K4%2BIH5zQJ8G%2F3qYUMh6%2Bs%3D","/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100005.wav","The focus- of this chapter is the American revolution%.DH AX0 / F OW1 . K AX0 S3 / AX1 V / DH IH1 S / CH AE1 P . T ER0 / IH1 Z / DH IY0 / AX0 . M EH1 . R IH0 . K AX0 N / R EH2 . V AX0 . L UW1 . SH AX0 N"],["mp3","https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100007.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=tHaBfAfUc8rIfSAOcZp%2F8a68TLM%3D","/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100007.wav","Can I go calling any time%?K AE1 N / AY13 / G OW1 / K AO1 . L IH0 NG / EH1 . N IY0 / T AY1 M"],["mp3","https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100004.wav?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=NfaAP7%2B%2BDLU0hQ5AkuwthPVJmj4%3D","/data/apps/damp/temp/ziptemp/APY220430001_demo1695809020325/APY220430001_demo/100004.wav","Is- it really take- time%?IH1 Z / IH1 T / R IY1 . AX0 . L IY0 / T EY1 K / T AY1 M3"]],"officialSummary":"Male audio data of American English. It is recorded by American English native speakers, with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.","dataexampl":"","datakeyword":["TTS","American English","Male"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Voice Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"speechSyn","BGimg":"brightSpot_audio","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"],"single":"no"}
20 Hours - American English Speech Synthesis Corpus-Male
TTS
American English
Male
Male audio data of American English. It is recorded by American English native speakers, with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Format
48,000Hz, 24bit, uncompressed wav, mono channel;
Recording environment
professional recording studio;
Recording content
general narrative sentences, interrogative sentences, etc;
Speaker
male, 20-30 years old, young and positive voice;
Device
microphone;
Language
American English;
Annotation
word and phoneme transcription, four-level prosodic boundary annotation;
Application scenarios
speech synthesis.
Sample
Audio
Look- at- the way they hear operas- and- see oil paintings%.L UH1 K3 / AE1 T / DH AX0 / W EY1 / DH EY1 / HH IY1 R / AA1 . P R AX0 Z / AX0 N D / S IY1 / OY1 L / P EY1 N . T IH0 NG Z
Audio
Was there some discussion- about- whether I should- speak%?W AX1 Z / DH EH1 R / S AH1 M / D IH0 . S K AH1 . SH AX0 N3 / AX0 . B AW1 T / W EH1 . DH ER0 / AY1 / SH UH1 D / S P IY1 K
Audio
The focus- of this chapter is the American revolution%.DH AX0 / F OW1 . K AX0 S3 / AX1 V / DH IH1 S / CH AE1 P . T ER0 / IH1 Z / DH IY0 / AX0 . M EH1 . R IH0 . K AX0 N / R EH2 . V AX0 . L UW1 . SH AX0 N
Audio
Can I go calling any time%?K AE1 N / AY13 / G OW1 / K AO1 . L IH0 NG / EH1 . N IY0 / T AY1 M
Audio
Is- it really take- time%?IH1 Z / IH1 T / R IY1 . AX0 . L IY0 / T EY1 K / T AY1 M3
Recommended Dataset
4 People - Northeastern dialect Average Tone Speech Synthesis Corpus
4 People - Northeastern dialect Average Tone Speech Synthesis Corpus. It is recorded by Northeast native. About 40% of the corpus contains words unique to Northeast China, the phonemes and tones are balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
2 People - Japanese Average Tone Speech Synthesis Corpus
2 People - Japanese Average Tone Speech Synthesis Corpus. It is recorded by rn native Japan, with authentic accent. Contains news and colloquial style general corpus,the phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
TTSJapaneseAverage Tone
10 Hours - Chaozhou Dialect Speech Synthesis Corpus - Female
10 Hours - Chaozhou Dialect Speech Synthesis Corpus - Female. It is recorded by Chaozhou-Shantou Pronunciation. the phonemes and tones are balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
Synthesis CorpusTTSFemaleGeneralChaozhouDialect
2 People - Mexican Spanish Average Tone Speech Synthesis Corpus
2 People - Mexican Spanish Average Tone Speech Synthesis Corpus. It is recorded by rn native Mexican, with authentic accent, Covering both customer service and general styles. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
TTSMexicanSpanishAverage Tone
2 People - Spanish Average Tone Speech Synthesis Corpus
2 People - Spanish Average Tone Speech Synthesis Corpus. It is recorded by rn native Spaniard, with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
TTSSpanishAverage Tone
2 People - New Zealand English Average Tone Speech Synthesis Corpus
2 People - New Zealand English Average Tone Speech Synthesis Corpus. It is recorded by rn native New Zealanders, with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
TTSNew Zealand EnglishAverage Tone
20 Hours - Sichuan Dialect Speech Synthesis Corpus - Female
20 Hours - Sichuan Dialect Speech Synthesis Corpus - Female. It is recorded by Chengdu Sichuan Pronunciation. the phonemes and tones are balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
Synthesis CorpusTTSFemaleGeneralSichuanDialect
10 People - British English Average Tone Speech Synthesis Corpus
10 People - British English Average Tone Speech Synthesis Corpus. It is recorded by British English native speakers, with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.