en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

NLU Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
34
Entity Identification
4
Dialogue Text
1
Intention Understanding
1
Others
2
Parallel Corpus
23

5.3M Pairs German-Chinese Parallel Corpus for NLP and MT Applications

5.3 million Chinese-German parallel sentence pairs stored in text format, covering multiple domains such as tourism, medical treatment, daily life, news, etc. The data desensitization and quality checking had been done. It can be used for machine translation, NLP research, and bilingual text analysis.
german parallel corpus chinese german sentence pairs dataset chinese german bilingual corpus chinese german NLP corpus chinese german text alignment dataset

English Intent Recognition Dataset – 84,516 Sentences with Slot Filling Annotations

This dataset contains 84,516 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, and smart home equipment, etc. It is applied to intent recognition, intent classification, and slot filling tasks.
intent recognition dataset intent classification dataset natural language understanding dataset NLU dataset English annotated intent dataset intent detection dataset dialogue intent dataset chatbot training dataset slot filling dataset

1,080,000 Groups – English-Russian Parallel Corpus Data

English and Russian parallel corpus, 1,080,000 groups in total; excluded political, porn, personal information and other sensitive vocabulary; it can be a base corpus for text-based data analysis, used in machine translation and other fields.
English-Russian parallel corpus

1,340,000 Groups – English-Korean Parallel Corpus Data

English and Korean parallel corpus, 1340,000 groups in total; excluded political, porn, personal information and other sensitive vocabulary; it can be a base corpus for text-based data analysis, used in machine translation and other fields.
English-Korean parallel corpus

Japanese-English Parallel Corpus Dataset – 380,000 Sentence Pairs for Machine Translation

This dataset contains 380,000 aligned Japanese-English sentence pairs. Sensitive content such as political, pornographic, and personal information has been excluded. It can be a base corpus for translation systems, cross-lingual information retrieval, and text-based language processing applications.
japanese english parallel corpus japanese english translation dataset bilingual corpus japanese english japanese english text dataset japanese english aligned corpus japanese english NLP dataset

Large-Scale Open Domain Intent Recognition Dataset – 687,694 Annotated Sentences

Annotation of 687,694 sentences generated by users in the mobile phone scene, covering to-do scenes, location scenes, and schedule scenes. The data set can be used for natural language understanding tasks.
intent recognition dataset intent classification dataset Chinese intent recognition data

Intent Classification Dataset – 47,811 Annotated English Sentences for Dialogue Systems

This dataset contains 47,811 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, smart home equipment, etc. it is applied to intent recognition, intent classification, and slot filling tasks.
slot filling dataset intent detection dataset intent recognition dataset intent classification dataset

English-Japanese Parallel Corpus – 850,000 Sentence Pairs for Machine Translation

This dataset contains 850,000 English-Japanese parallel sentences stored in TXT format. It covers multiple fields such as tourism, medical treatment, daily life, news, etc. average English sentence 23 words. The data desensitization and quality checking had been done. It can be used as a fundamental dataset for machine translation, bilingual NLP tasks, and other text processing applications.
English Japanese parallel corpus English Japanese translation dataset English Japanese bilingual corpus English Japanese parallel dataset English Japanese text dataset English Japanese MT dataset

Chinese-English Parallel Corpus Dataset (80,120,000 Sentence Pairs) – Translation & NLP

This dataset contains 80 million Chinese-English parallel sentences, covering domains such as travel, medicine, daily conversation, and TV scripts. It is stored in txt format, cleaned, desensitized, and quality-checked. It can be used as a fundamental dataset for machine translation, bilingual NLP tasks, and other text processing applications.
Chinese English parallel corpus Chinese English translation dataset Chinese English machine translation data Chinese English bilingual corpus Chinese English parallel dataset Chinese English text dataset

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
f278579e-0dbb-4349-abe4-8982fea64cce