en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

8 Hours – Cantonese Speech Dataset for TTS (Hong Kong)

Cantonese speech dataset
Hong Kong Cantonese speech corpus
Cantonese text-to-speech dataset
Cantonese voice dataset for AI
native Cantonese speech recordings
Cantonese TTS dataset
Hong Kong accent speech dataset

This dataset features recordings from 4 native Hong Kong Cantonese speakers. The corpus contain educational, game and general colloquial content. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It precisely matches with the research and development needs of the speech synthesis.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Format
48,000Hz, 24bit, uncompressed wav, mono channel;
Recording environment
professional recording studio;
Recording content
contains educational, game and general colloquial content;
Speaker
professional voice actor, two male and two female, 2 hours per person;
Annotation
word and phoneme transcription, prosodic boundary annotation;
Device
microphone;
Language
Hong Kong Cantonese;
Application scenarios
speech synthesis.
Sample Sample
  • Audio

    佢哋#1產生喺#1新文化#1運動#1之後#4。 keoi5 dei6 caan2 sang1 hai2 san1 man4 faa3 wan6 dung6 zi1 hau6

  • Audio

    好多#1遊戲#1都#1只係#1換個#1皮膚#1做#1副本#4。 hou2 do1 jau4 hei3 dou1 zi2 hai6 wun6 go3 pei4 fu1 zou6 fu3 bun2

  • Audio

    事物#1相應嘅#1心理#1需要#1而#1產生#4。 si6 mat6 soeng1 jing3 ge3 sam1 lei5 seoi1 jiu3 ji4 caan2 sang1

  • Audio

    你#1日頭#1又#1可以#2夜晚#1又#1可以#4。 lei5 jat6 tau2 jau6 ho2 ji3 je6 maan5 jau6 ho2 ji3

Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

What languages and voice characteristics are covered by Nexdata’s speech synthesis datasets?

Nexdata offers speech synthesis datasets covering a broad range of languages, dialects, accents, and voice types, supported by extensive global language resources. Our datasets include diverse speakers, speaking styles, emotions, and recording scenarios to support natural and expressive Text-to-Speech (TTS) model development.

Can Nexdata customize speech synthesis datasets for specific languages or requirements?

Yes. If our off-the-shelf TTS datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, annotation, and quality control services. We can customize datasets based on target languages or dialects, speaker profiles, voice characteristics, emotions, speaking styles, recording environments, and data volume.

How does Nexdata ensure the quality and scalability of speech synthesis datasets?

Nexdata applies multi-stage quality control throughout voice data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual TTS projects and specialized datasets for specific voices, accents, emotions, or speech scenarios.

37a6acf6-821b-4919-8bb6-3495cb4ebf90

3516e512-68d6-4905-b4f7-fd72c6cfa2a8