en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

NLU Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
34
Entity Identification
4
Dialogue Text
1
Intention Understanding
1
Others
2
Parallel Corpus
23

84,516 Sentences - English Intention Annotation Data in Interactive Scenes

84,516 Sentences - English Intention Annotation Data in Interactive Scenes, annotated with intent classes, including slot and slot value information; the intent field includes music, weather, date, schedule, home equipment, etc.; it is applied to intent recognition research and related fields.
English Intent-type Intention

75 Dictionaries of Different Chinese Fields

75 Chinese domain dictionaries, including data for a certain year and covering a wide range of content. Each line in the data file includes a term and its Chinese pinyin, and the terms are sorted alphabetically. This data set can be used for tasks such as natural language understanding, knowledge base building, etc..
Chinese domain dictionary data text data NLU data Entity Identification data

7,290,000 Groups -Chinese -Vietnamese Parallel Corpus Data

7.29 Million Pairs of Sentences - Chinese-Vietnamese Parallel Corpus Data be stored in text format. It covers multiple fields such as tourism, medical treatment, daily life, news, etc. The data desensitization and quality checking had been done. It can be used as a basic corpus for text data analysis in fields such as machine translation.
Chinese Vietnamese Chinese-Vietnamese Parallel Corpus

8,178 Chinese Social Comments Events Annotation Data

8,178 Chinese social comments annotated data. The contents are hot news in 2013. Each piece of news contains one or more events and is annotated with time, theme, cause, procedure and result. The data is stored in xml and can be used for natural language understanding.
Social commentary Event Annotation

10,000 Chinese News Events Annotation Data

10,000 Chinese news event annotated data. The contents are hot news in 2013. Each piece of news contains one or more events. Each event is annotated. The data is stored in xml and can be used for natural language understanding.
News Events Annotation

12,820,000 Groups - Chinese-Korean Parallel Corpus Data

12,820,000 sets of parallel translation corpus between China and Korea, which are stored in txt files. It covers many fields including spoken language, traveling, news, and finance. Data cleaning, desensitization, and quality inspection have been carried out. It can be used as the basic corpus database in the text data files as well as used in machine translation.
Chinese Korean Chinese-Korean Parallel Corpus

980,000 Groups - Chinese-Urdu Parallel Corpus Data

980,000 sets of Chinese and Urdu language parallel translation corpus, data storage format is txt document. Data cleaning, desensitization, and quality inspection have been carried out, which can be used as a basic corpus for text data analysis and in fields such as machine translation.
Chinese Urdu Chinese-Urdu Parallel Corpus

1,990,000 Groups - Chinese-Czech Parallel Corpus Data

1,990,000 sets of Chinese and Czech language parallel translation corpus, data storage format is txt document. Data cleaning, desensitization, and quality inspection have been carried out, which can be used as a basic corpus for text data analysis and in fields such as machine translation.
Chinese Czech Parallel

1,980,000 Groups - Chinese-Polish Parallel Corpus Data

1,980,000 sets of Chinese and Polish language parallel translation corpus, data storage format is txt document. Data cleaning, desensitization, and quality inspection have been carried out, which can be used as a basic corpus for text data analysis and in fields such as machine translation.
Chinese Polish Parallel

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
e7b2d4c6-0b3c-413f-9239-c433ab3de3fa