en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

2 People - Vietnamese‌ Natural Conversation Average Tone Speech Synthesis Corpus

Vietnamese
Natural Conversation
Freetalk
Average Tone
TTS
Emotion
Multi-level Emotion
Paralanguage
Interjection

Vietnamese Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Format
48kHz, 24bit, uncompressed wav, mono channel
Recording Environment
Professional recording studio
Recording Content
Contains following categories: topic(Extemporaneous speaking on a given topic), multi-level emotion, single-level emotion, paralanguage
Personnel
professional voice actors; 1 male, 1 female
Annotation Features
Text annotation, emotion annotation, paralinguistic annotation
Equipment
Professional recording devices and software
Language
Vietnamese‌
Application Scenario
Speech Synthesis
Sample Sample
  • Audio

    User:Âm thanh có phải càng ngày càng gần hơn không? ||| [fear3] Hình như là vậy. Bọn mình phải làm sao bây giờ? Phòng bên cạnh hình như cũng không có người.

  • Audio

    User:Cậu hẹn mà đi muộn ba mươi phút rồi đó. ||| Tớ xin lỗi nhiều nha, tại nay tắc đường quá nên mới tới trễ vậy. Đã vậy sáng nay lúc ra khỏi nhà mãi mới tìm thấy chìa khoá xe cậu ạ.

  • Audio

    User:Cậu nghe tin gì chưa? ||| [surprise1] Nghe rồi. Đức Anh sắp cưới phải không? Ôi trời ơi biết tin này tớ giật mình luôn.

  • Audio

    User:Cậu quên kéo khóa quần kìa. ||| Hả, thật hả? Ôi ngại chết đi được. Cái này hình như bị hỏng cậu ạ, không phải do tớ quên đâu. Tớ chạy về nhà lấy quần khác thay luôn đây, ngại quá.

  • Audio

    User:Người đó vừa hỏi số của cậu đúng không? ||| Ừ, ngại chết đi được. Bạn đấy cũng đẹp trai cơ. Tự dưng chạy ra hỏi số điện thoại tớ, tớ cũng không biết trả lời sao nữa. Tớ ngại quá nên cứ lúng túng mãi, không biết nói gì luôn. Chị đồng nghiệp thì cứ đứng nhìn tớ cười, tớ ngại quá nên cứ đỏ mặt mãi. Tớ cũng không biết là có nên cho số điện thoại hay không nữa.

Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

What languages and voice characteristics are covered by Nexdata’s speech synthesis datasets?

Nexdata offers speech synthesis datasets covering a broad range of languages, dialects, accents, and voice types, supported by extensive global language resources. Our datasets include diverse speakers, speaking styles, emotions, and recording scenarios to support natural and expressive Text-to-Speech (TTS) model development.

Can Nexdata customize speech synthesis datasets for specific languages or requirements?

Yes. If our off-the-shelf TTS datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, annotation, and quality control services. We can customize datasets based on target languages or dialects, speaker profiles, voice characteristics, emotions, speaking styles, recording environments, and data volume.

How does Nexdata ensure the quality and scalability of speech synthesis datasets?

Nexdata applies multi-stage quality control throughout voice data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual TTS projects and specialized datasets for specific voices, accents, emotions, or speech scenarios.

37300759-67ba-44c5-9840-aaf02c4b4885

6b78d8b4-84d1-4dbc-b8c7-129c01fab340