en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

593 Hours - English(China) Scripted Monologue Smartphone speech dataset

Chinese speak English
English voice data
and mobile phones collect voice data.china
chinaman
oriental
taiwanese
byzantine
chink
chinese
woman
formosan
asian
chinawoman
japanese
korean
sinaean
celestial
chine
chino
chugoku
non-chinese
ovary
pekin
acupuncture
airframe
airframes
all-china
beijing
patois
vernacular
language
lingo
speech
idiom
argot
tongue
accent
jargon
slang
parlance
cant
patter
brogue
provincialism
localism
regional
language
locution
terminology
vocabulary
local
speech
regionalism
pidgin
regionalisms
colloquialism
phraseology
idioms
mother
tongue
talk
pronunciation
local
language
creole
langue
lingua
franca
phrasing
tongues
idiolect
lexicon
vernacularism
colloquial
word
brogues
dialectal
jive
talk
languages
lingua
localisms
wording
accents
business
language

English(China) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering 100,000 common expressions. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(3,691 Chinese, covering domestic dialect zones like Jiangsu, Shandong, Beijing, He'nan, and meets the specific accents of Chinese speaking English), geographicly speaking, enhancing model performance in real and complex tasks.nQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Format
16kHz, 16bit, uncompressed wav, mono channel;
Recording condition
Low background noise (indoor), without echo;
Content category
100,000 common expressions;
Recording device
Android smartphone;
Speaker
3,691 Chinese, 34% male and 66% female;
Country
China(CHN);
Language
English;
Features of annotation
Transcription text;
Accuracy Rate
Sentence Accuracy Rate (SAR) 95%
Sample Sample
  • Audio

    Her cheeks had fallen in,making her look old.

  • Audio

    No milk. I'm slimming.

  • Audio

    We're focused on small things: Do I have my pierce?

  • Audio

    No one could know why he did like that.

  • Audio

    The bark scaled off the tree.

Recommended DatasetsRecommended Dataset
2,028 Hours - Mandarin(China) Scripted Monologue Smartphone speech dataset

Mandarin(China) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.rnQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

mandarin speech data Scripted Monologue speech data chinese speech data
11,010 People - Mandarin(China) Digital Smartphone speech dataset

Mandarin(China) Digital Smartphone speech dataset, each speaker reads 30 sentences of 4 -8 digit number.Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Mandarin Digital voice print
18 Hours - English(Brazil) Scripted Monologue Smartphone speech dataset

English(Brazil) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(55 people in total), geographicly speaking, enhancing model performance in real and complex tasks.rnQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Accent English Brazil English
207 Hours - English(Japan) Scripted Monologue Smartphone speech dataset

English(Japan) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(464 people in total), geographicly speaking, enhancing model performance in real and complex tasks.nQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Accent English Japanese Japan English
207 Hours - English(Canada) Scripted Monologue Smartphone speech dataset

English(Canada) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(466 people in total), geographicly speaking, enhancing model performance in real and complex tasks.nQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Canada English Accent English asr datasets
199 Hours - English(Australia) Scripted Monologue Smartphone speech dataset

English(Australia) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(402 people in total), geographicly speaking, enhancing model performance in real and complex tasks.rnQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Australia causal speech data Australia causal data Australia causal dataset Australia causal conversation Australia causal conversation data Australia causal conversation dataset Australia causal chat data
201 Hours - English(Singapore) Scripted Monologue Smartphone speech dataset

English(Singapore) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(452 people in total), geographicly speaking, enhancing model performance in real and complex tasks.rnQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

English speech data singaporean speech dataset singapore english speech data
198 Hours - English(Malaysia) Scripted Monologue Smartphone speech dataset

English(Malaysia) Scripted Monologue Smartphone speech dataset, collected from monologue based on given scripts, covering generic domain, human-machine interaction, smart home command and control, in-car command and control, numbers and other domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(423 people in total), geographicly speaking, enhancing model performance in real and complex tasks.rnQuality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Accent English Malaysia English
Tell Us Your Special Needs

By submitting, I agree to the Privacy Protection

af6af7ac-860a-44d2-b9d8-924974d50f90

52f9d793-41cc-4908-af39-09e85b6d8881