en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

High-Quality Training Datasets

Boost the performance of your AI models with our high-quality, ready-to-use training datasets.

Language

All

Data Type

All

2 People - French Natural Conversation Average Tone Speech Synthesis Corpus

French Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
French Natural Conversation Average Tone TTS Emotion Multi-level Emotion Paralanguage Interjection Freetalk

2 People - Indonesian Natural Conversation Average Tone Speech Synthesis Corpus

Indonesian Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of modal particles). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Indonesian Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Interjection Paralanguage

2 People - Saudi Arabian Arabic Natural Conversation Average Tone Speech Synthesis Corpus

Saudi Arabian Arabic Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Saudi Arabian Arabic Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Paralanguage Interjection

2 People - Modern Standard Arabic Natural Conversation Average Tone Speech Synthesis Corpus

Modern Standard Arabic Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Modern Standard Arabic Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Paralanguage interjection

1,042 Segments 6-camera Egocentric Embodied AI Dataset

This dataset contains 1,042 egocentric video segments (approximately 35 seconds each) collected across 34 locations and 6 real-world environments, including homes, offices, and retail scenarios. Powered by self-developed VSLAM system, achieving millimeter-level positioning, multi-sensor hard-triggered synchronization (≤1ms), and a high frame rate of 60fps for RGB. It includes multi-view videos, calibrations, point clouds, SLAM trajectories, and gesture recognition results in standard formats. Designed for embodied AI, spatial perception, and 3D reconstruction, it offers high precision, diverse scenarios, and out-of-the-box usability, making it ideal for training robust perception-action models.
embodied ai dataset robot learning dataset robotics training data egocentric dataset robot perception dataset multimodal robotics dataset SLAM dataset

2 People - Japanese Natural Conversation Average Tone Speech Synthesis Corpus

Japanese Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of modal particles). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Japanese Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Paralanguage Modal Particle

2 People - Italian Natural Conversation Average Tone Speech Synthesis Corpus

Italian Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Italian Natural Conversation Freetalk TTS Emotion Multi-level Emotion Paralanguage Modal Particle Average Tone Interjection

2 People - Brazilian Portuguese Natural Conversation Average Tone Speech Synthesis Corpus

Brazilian Portuguese Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Brazilian Portuguese Natural Conversation Freetalk Average Tone TTS Emotion Multi-level Emotion Paralanguage Interjection

2 People - German Natural Conversation Average Tone Speech Synthesis Corpus

German Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
German Average Tone Natural Conversation Freetalk Emotion TTS Multi-level Emotion Paralanguage Interjection

2 People - Indian English Natural Conversation Average Tone Speech Synthesis Corpus

Indian English Natural Conversational Speech Synthesis Corpus. Recorded by two native voice talents (one male, one female), this corpus comprises two modules: free dialogue on specified topics and scripted emotional/paralinguistic text performances. It covers single-emotion, multi-emotion, and paralinguistic expressions (including a variety of interjections). Professional linguists have provided three-tier annotation—on text content, emotion labels, and paralinguistic phenomena. The total duration is approximately 11 hours, with audio specifications of 48 kHz sampling rate, 24‑bit depth, WAV format, mono channel. This dataset is designed to fully meet the diverse requirements of TTS research and development in terms of naturalness, expressiveness, and precise annotation.
Indian English Natural Conversation Average Tone Freetalk TTS Emotion Multi-level Emotion Paralanguage Interjection

INTERSPEECH 2026 MLC-SLM Challenge Dataset

The INTERSPEECH 2026 MLC-SLM Challenge Dataset, curated by Datatang, is derived from fifteen proprietary conversational speech corpora. Distinguished by exceptional annotation accuracy and operational reliability, this dataset is engineered to address critical challenges in multilingual automatic speech recognition (ASR) and long-context comprehension. It meticulously replicates real-world complexities including spontaneous interruptions and speaker overlaps, thereby providing robust training resources for developing world-ready ASR systems. All data collection and processing strictly comply with international privacy regulations including GDPR, CCPA and PIPL, with rigorous protocols ensuring participant anonymity and ethical data usage throughout the lifecycle.
INTERSPEECH MLC-SLM

Agent Trajectory Dataset for Tool-Use Training and AI Agent Evaluation

This dataset covers office-based scenarios such as in-depth searches, data analysis, and industry research, encompassing complete multi-turn reasoning trajectories and tool-calling chains. It is designed to support the analysis of agent planning capabilities, research into tool selection strategies, and quality assessment, providing a structured benchmark for agent training and evaluation.
ai agent dataset agent training dataset agent trajectory dataset llm agent dataset tool use dataset agent evaluation datase
. . .
loading

loading

b3c410cb-4409-4eaf-8488-d2db06ca44ac