en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

NLU Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
34
Entity Identification
4
Dialogue Text
1
Intention Understanding
1
Others
2
Parallel Corpus
23

5.3M Pairs German-Chinese Parallel Corpus for NLP and MT Applications

5.3 million Chinese-German parallel sentence pairs stored in text format, covering multiple domains such as tourism, medical treatment, daily life, news, etc. The data desensitization and quality checking had been done. It can be used for machine translation, NLP research, and bilingual text analysis.
german parallel corpus chinese german sentence pairs dataset chinese german bilingual corpus chinese german NLP corpus chinese german text alignment dataset

English Intent Recognition Dataset – 84,516 Sentences with Slot Filling Annotations

This dataset contains 84,516 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, and smart home equipment, etc. It is applied to intent recognition, intent classification, and slot filling tasks.
intent recognition dataset intent classification dataset natural language understanding dataset NLU dataset English annotated intent dataset intent detection dataset dialogue intent dataset chatbot training dataset slot filling dataset

1,080,000 Groups – English-Russian Parallel Corpus Data

English and Russian parallel corpus, 1,080,000 groups in total; excluded political, porn, personal information and other sensitive vocabulary; it can be a base corpus for text-based data analysis, used in machine translation and other fields.
English-Russian parallel corpus

1,340,000 Groups – English-Korean Parallel Corpus Data

English and Korean parallel corpus, 1340,000 groups in total; excluded political, porn, personal information and other sensitive vocabulary; it can be a base corpus for text-based data analysis, used in machine translation and other fields.
English-Korean parallel corpus

Japanese-English Parallel Corpus Dataset – 380,000 Sentence Pairs for Machine Translation

This dataset contains 380,000 aligned Japanese-English sentence pairs. Sensitive content such as political, pornographic, and personal information has been excluded. It can be a base corpus for translation systems, cross-lingual information retrieval, and text-based language processing applications.
japanese english parallel corpus japanese english translation dataset bilingual corpus japanese english japanese english text dataset japanese english aligned corpus japanese english NLP dataset

Large-Scale Open Domain Intent Recognition Dataset – 687,694 Annotated Sentences

Annotation of 687,694 sentences generated by users in the mobile phone scene, covering to-do scenes, location scenes, and schedule scenes. The data set can be used for natural language understanding tasks.
intent recognition dataset intent classification dataset Chinese intent recognition data

Intent Classification Dataset – 47,811 Annotated English Sentences for Dialogue Systems

This dataset contains 47,811 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, smart home equipment, etc. it is applied to intent recognition, intent classification, and slot filling tasks.
slot filling dataset intent detection dataset intent recognition dataset intent classification dataset

1,990,000 Groups - Chinese-Czech Parallel Corpus Data

1,990,000 sets of Chinese and Czech language parallel translation corpus, data storage format is txt document. Data cleaning, desensitization, and quality inspection have been carried out, which can be used as a basic corpus for text data analysis and in fields such as machine translation.
Chinese Czech Parallel

10 Million Traditional Chinese Oral Message Data

Traditional Chinese SMS corpus, 10 million in total, real traditional Chinese spoken language text data; only contains text messages; the content is stored in txt format; the data set can be used for natural language understanding and related tasks.
Traditional Chinese SMS corpus traditional Chinese SMS data traditional Chinese SMS collection traditional Chinese corpus data

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
59ba899c-823f-41e2-b049-3647f8b46492