en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

NLU Datasets

Nexdata delivers multilingual NLU datasets for intent detection, entity recognition, and intelligent language understanding applications.

Type

All
26
Intention Understanding
3
Parallel Corpus
23

5.3M Pairs German-Chinese Parallel Corpus for NLP and MT Applications

5.3 million Chinese-German parallel sentence pairs stored in text format, covering multiple domains such as tourism, medical treatment, daily life, news, etc. The data desensitization and quality checking had been done. It can be used for machine translation, NLP research, and bilingual text analysis.
german parallel corpus chinese german sentence pairs dataset chinese german bilingual corpus chinese german NLP corpus chinese german text alignment dataset

English Intent Recognition Dataset – 84,516 Sentences with Slot Filling Annotations

This dataset contains 84,516 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, and smart home equipment, etc. It is applied to intent recognition, intent classification, and slot filling tasks.
intent recognition dataset intent classification dataset natural language understanding dataset NLU dataset English annotated intent dataset intent detection dataset dialogue intent dataset chatbot training dataset slot filling dataset

1.08 Million English Russian Parallel Corpus Dataset for Machine Translation

This dataset contains 1.08 million English-Russian sentence pairs, the corpus consists of aligned English and Russian text data covering diverse topics and general language usage scenarios. Sensitive content, including political, adult, and personally identifiable information (PII), has been filtered and removed. it can be a base corpus for machine translation, natural language processing, and multilingual AI model development.
english russian parallel corpus english russian translation dataset english russian bilingual dataset parallel corpus dataset machine translation dataset

1.34 Million English Korean Parallel Corpus Dataset for Machine Translation

This dataset contains 1.34 million English-Korean sentence pairs, providing high-quality bilingual text resources for machine translation, natural language processing, and multilingual AI model training. The corpus consists of aligned English and Korean text data covering diverse topics and general language usage scenarios. Sensitive content, including political, adult, and personally identifiable information (PII), has been filtered and removed. It can be a base corpus for text-based data analysis, used in machine translation and other fields.
english korean parallel corpus english korean translation dataset korean english translation dataset korean bilingual dataset machine translation dataset parallel corpus dataset

Japanese-English Parallel Corpus Dataset – 380,000 Sentence Pairs for Machine Translation

This dataset contains 380,000 aligned Japanese-English sentence pairs. Sensitive content such as political, pornographic, and personal information has been excluded. It can be a base corpus for translation systems, cross-lingual information retrieval, and text-based language processing applications.
japanese english parallel corpus japanese english translation dataset bilingual corpus japanese english japanese english text dataset japanese english aligned corpus japanese english NLP dataset

Intent Classification Dataset – 47,811 Annotated English Sentences for Dialogue Systems

This dataset contains 47,811 English sentences annotated with intent classes, slot labels, and slot values. The intent field includes music, weather, date, schedule, smart home equipment, etc. it is applied to intent recognition, intent classification, and slot filling tasks.
slot filling dataset intent detection dataset intent recognition dataset intent classification dataset

750K Chinese-Burmese Sentence Pairs – Machine Translation Dataset

This dataset contains 750,000 Chinese-Burmese parallel sentence pairs stored in TXT format. The corpus covers multiple domains, including tourism, healthcare, daily life, news, and other topics. The data has undergone cleaning, anonymization, and quality inspection to improve data quality and protect sensitive information. It can be used as a basic corpus for text data analysis in fields such as machine translation.
Chinese Burmese parallel corpus Chinese Burmese translation dataset Chinese Burmese parallel dataset Chinese Burmese sentence pairs Chinese Burmese translation corpus

7.29M Chinese-Vietnamese Sentence Pairs – Machine Translation Dataset

This dataset contains 7.29 million Chinese-Vietnamese parallel sentence pairs stored in TXT format. The corpus covers multiple domains, including tourism, healthcare, daily life, news, and other topics. The data has undergone cleaning, anonymization, and quality inspection to improve data quality and protect sensitive information. It can be used as a basic corpus for text data analysis in fields such as machine translation.
Chinese Vietnamese parallel corpus Chinese Vietnamese translation dataset Chinese Vietnamese parallel dataset Chinese Vietnamese sentence pairs Chinese Vietnamese translation corpus

9.83 Million Chinese Japanese Bilingual Corpus for NLP and LLM Training

This dataset contains 9.83 million Chinese-Japanese sentence pairs stored in TXT format. The corpus covers multiple domains, including general topics, information technology, news, patents, and other specialized fields. Each sentence pair has undergone data anonymization and quality assurance processes to ensure usability for AI model development. It can be used as a basic corpus for text data analysis in fields such as machine translation.
chinese japanese parallel corpus chinese japanese translation dataset chinese japanese bilingual dataset parallel corpus dataset machine translation dataset

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
548b1eaa-1a5a-4819-8289-b8ef247a2d76