en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

Vietnamese TTS Dataset - 2 Native Speakers Speech Synthesis Corpus

vietnamese tts dataset
vietnamese speech synthesis corpus
vietnamese speech corpus
vietnamese voice dataset
vietnamese text to speech dataset

This Vietnamese speech synthesis corpus is recorded by 2 native Vietnamese speakers with authentic Vietnamese accents. It features balanced phoneme coverage and was annotated with the participation of professional phoneticians, enabling it to precisely meet the R&D requirements for speech synthesis.

Paid Datasets
This is a paid dataset licensed for commercial use. Ready-made datasets are available for immediate integration into AI projects.
SpecificationsSpecifications
Format
48,000Hz, 24bit, uncompressed wav, mono channel;
Recording environment
professional recording studio;
Recording content
general corpus;
Speaker
Vietnamese, 1 male and 1 female, 10 hours per person;
Annotation
word and phoneme transcription, prosodic boundary annotation, phoneme boundary annotation;
Device
microphone;
Language
Vietnamese;
Application scenarios
speech synthesis.
Sample Sample
  • Audio

    Nhớ hồi xưa đi học trốn tiết ra sân tập%, bị bác bảo vệ đuổi chạy toát mồ hôi hột%, vui thật đấy%. {ɲ əː3} {h o2 i} {s ɯə1} {ɗ i1} {h ɔ6 k} {c o3 n} {t ie3 t} {z aː1} {s ə1 n} {t ə6 p}, {ɓ i6} {ɓ aː3 k} {ɓ aː4 u} {v e6} {ɗ uo4 i} {c a6 i} {t u aː3 t} {m o2} {h o1 i} {h o6 t}, {v u1 i} {tʰ ə6 t} {ɗ ə3 i}.

  • Audio

    Bà bây giờ cũng sành điệu ghê%. Thế chốt nhé%, ba giờ chiều gặp nhau ở chân đê%. {ɓ aː2} {ɓ ə1 i} {z əː2} {k u5 ŋ} {s aː2 ɲ} {ɗ ie6 u} {ɣ e1}. {tʰ e3} {c o3 t} {ɲ ɛ3}, {ɓ aː1} {z əː2} {c ie2 u} {ɣ a6 p} {ɲ a1 u} {əː4} {c ə1 n} {ɗ e1}.

  • Audio

    Nước sông chuyển sang màu lạ và bốc mùi hôi% là dấu hiệu rõ ràng% của ô nhiễm hữu cơ nặng%. {n ɯə3 k} {s o1 ŋ} {c u ie4 n} {s aː1 ŋ} {m a2 u} {l aː6} {v aː2} {ɓ o3 k} {m u2 i} {h o1 i} {l aː2} {z ə3 u} {h ie6 u} {z ɔ5} {z aː2 ŋ} {k uə4} {o1} {ɲ ie5 m} {h ɯ5 u} {k əː1} {n a6 ŋ}.

  • Audio

    Đúng rồi%. Mở ra thấy toàn là những kỷ vật của cha mẹ và chúng mình hồi còn bé%. {ɗ u3 ŋ} {z o2 i}. {m əː4} {z aː1} {tʰ ə3 i} {t u aː2 n} {l aː2} {ɲ ɯ5 ŋ} {k i4} {v ə6 t} {k uə4} {c aː1} {m ɛ6} {v aː2} {c u3 ŋ} {m i2 ɲ} {h o2 i} {k ɔ2 n} {ɓ ɛ3}.

  • Audio

    Văn học kinh điển/ vẫn luôn giữ nguyên giá trị nhân văn sâu sắc% bởi phản ánh chân thực khát vọng/ và đấu tranh của con người/ qua nhiều thế kỷ%. {v a1 n} {h ɔ6 k} {k i1 ɲ} {ɗ ie4 n} {v ə5 n} {l uo1 n} {z ɯ5} {ŋ u ie1 n} {z aː3} {c i6} {ɲ ə1 n} {v a1 n} {s ə1 u} {s a3 k} {ɓ əː4 i} {f aː4 n} {aː3 ɲ} {c ə1 n} {tʰ ɯ6 k} {x aː3 t} {v ɔ6 ŋ} {v aː2} {ɗ ə3 u} {c aː1 ɲ} {k uə4} {k ɔ1 n} {ŋ ɯə2 i} {k u aː1} {ɲ ie2 u} {tʰ e3} {k i4}.

Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

Dataset FAQs

What languages and voice characteristics are covered by Nexdata’s speech synthesis datasets?

Nexdata offers speech synthesis datasets covering a broad range of languages, dialects, accents, and voice types, supported by extensive global language resources. Our datasets include diverse speakers, speaking styles, emotions, and recording scenarios to support natural and expressive Text-to-Speech (TTS) model development.

Can Nexdata customize speech synthesis datasets for specific languages or requirements?

Yes. If our off-the-shelf TTS datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, annotation, and quality control services. We can customize datasets based on target languages or dialects, speaker profiles, voice characteristics, emotions, speaking styles, recording environments, and data volume.

How does Nexdata ensure the quality and scalability of speech synthesis datasets?

Nexdata applies multi-stage quality control throughout voice data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual TTS projects and specialized datasets for specific voices, accents, emotions, or speech scenarios.

2ac2d2dc-41f6-4c17-b565-501933c950eb

f629c6bb-612f-4227-9d6f-097e23493b85