en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Speech Synthesis Datasets

Nexdata provides high-quality TTS datasets covering diverse languages, speaking styles, emotions, and recording scenarios for advanced TTS model training.

Voice Type

All
40
Average Tone
32
Customer Service
2
Emotion
18
Female
8
Male
5
Tone Quality
2

Language

All
40
Chinese Dialects
2
English
11
Japanese
3
Mandarin
1
Others
25

20 Hours - American English Male Voice TTS Dataset

This dataset contains 20 hours of American English male voice recordings. It is recorded by Americans (native English speakers) with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It is suitable for text-to-speech (TTS) model training, phoneme recognition research, and AI voice development.
TTS english dataset speech synthesis dataset TTS male voice dataset male voice dataset for tts American English speech synthesis dataset

19.46 Hours - American English Female Voice TTS Dataset

This dataset contains 19.46 hours of American English female voice recordings. It is recorded by American (native English speaker) with authentic accent and clear, sweet tone. The phoneme coverage is balanced. Professional phoneticians participate in the annotation. It is suitable for text-to-speech (TTS) model training, phoneme recognition, and AI voice development requiring natural-sounding female speech.
American English speech synthesis dataset female voice dataset for TTS American English female voice corpus speech synthesis training data female TTS dataset American English female speaker speech synthesis dataset TTS english dataset

10.4 Hours – Japanese Female Voice TTS Dataset

This dataset contains 10.4 hours of Japanese female voice recordings. It is recorded by Japanese native speaker with an authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. This corpus is ideal for tasks such as Japanese text-to-speech (TTS) training, speech synthesis research, and AI voice model development.
Japanese speech synthesis dataset Japanese tts dataset Japanese text-to-speech dataset female female japanese tts dataset

2 Speakers – Australian English TTS Dataset (Native Accent)

This dataset features recordings from 2 native Australian English speakers with authentic accents. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
Australian English TTS dataset Australian speech dataset for AI Australian accent speech dataset Australian text to speech voices multi-speaker Australian English dataset Australian English phoneme balanced dataset

8 Hours – Spanish TTS Dataset with Native Castilian Accent

This dataset includes recordings from 2 native Spanish speakers with authentic Castilian accents. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
Spanish speech dataset for TTS Spanish text to speech dataset Spanish voice dataset for AI models native Spanish accent dataset Castilian Spanish TTS dataset Spanish speech synthesis dataset

12 Hours – Italian TTS Dataset with Native Accent

This dataset includes recordings from 3 native Italian speakers with authentic accents. Covering both customer service and general speaking styles. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.
Italian speech dataset for TTS Italian text to speech dataset Italian voice dataset for AI Italian accent speech dataset multi-speaker Italian TTS dataset Italian TTS dataset

10 Speakers – British English TTS Dataset with Authentic Accent

This dataset contains recordings from 10 native British English speakers with an authentic UK accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the text-to-speech (TTS) systems and AI voice synthesis models.
British English speech synthesis dataset British English voice dataset for TTS British accent speech corpus UK English speech dataset female male natural British English voice dataset British English tts dataset

8 Hours - Canadian French TTS Dataset (Native Accent)

This dataset contains recordings from 2 native Canadian French speakers with authentic accents. It is ideal for researchers and developers seeking natural Canadian French voices.
Canadian French TTS dataset Canadian French speech dataset for AI Canadian French accent speech corpus Canadian French text to speech voices Canadian French speech dataset

2 People - Indonesian Natural Conversation Average Tone Speech Synthesis Corpus

Indonesian Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of modal particles). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Indonesian Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Interjection Paralanguage

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
427cbd3a-c61e-412e-9833-3c2c104d9fce