en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

204 Hours Filipino English Speech Dataset with Real-world Conversational Audio

filipino english speech dataset
philippine english speech dataset
customer service speech dataset
english ASR dataset
voice assistant training data
speech recognition training data

This dataset contains 204 hours of spontaneous conversational English speech collected from Filipino speakers using smartphone devices. The recordings are based on topic-guided conversations covering diverse real-world scenarios. Each audio sample is transcribed with corresponding text content, timestamps, speaker ID, gender, and other metadata attributes. This dataset collects data from around 400 native English-speaking Filipinos, capturing natural conversational patterns, pronunciation variations, and regional English features to support the development of robust speech AI models and improve their performance in real-world and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Format
16kHz, 16bit, uncompressed wav, mono channel
Content category
Dialogue based on given topics
Recording condition
Low background noise (indoor)
Recording device
Android smartphone, iPhone
Country
Philippine(PHL)
Language(Region) Code
en-PH
Language
English
Speaker
304 native speakers in total, 42% male and 58% female
Features of annotation
Transcription text, timestamp, speaker ID, gender, noise
Accuracy rate
Sentence accuracy rate(SAR) 95%
Sample Sample
  • Audio

    I'm fine, thank you for asking.

  • Audio

    I've never thought that you are into Islam.

  • Audio

    Today is Sunday, do you have any plan on going anywhere?

  • Audio

    Oh yeah, I will actually on my way going to our church.

  • Audio

    Oh actually I go to the mosque at Friday.

Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

b39aef85-0a72-47c7-b6d7-76756084803c

56fd643f-ebd0-4ce5-be52-3bda1f0aefd2