[{"@type":"PropertyValue","name":"Format","value":"TXT"},{"@type":"PropertyValue","name":"Data content","value":"Chinese-Japanese parallel corpus"},{"@type":"PropertyValue","name":"Data size","value":"9.83 million pairs of Chinese-Japanese Parallel Corpus Data."},{"@type":"PropertyValue","name":"Language","value":"Chinese, Japanese"},{"@type":"PropertyValue","name":"Applications","value":"machine translation"},{"@type":"PropertyValue","name":"Accuracy rate","value":"90%"}]
{"id":1069,"datatype":"1","titleimg":"https://res.datatang.com/asset/productNew/APY200214001.png?Expires=2007353678&OSSAccessKeyId=LTAI5tQwXnJZbubgVfVa1ep9&Signature=AiMVpzHU8YYk4SEGm8N2lhg5ryA%3D","type1":"183","type1str":null,"type2":"185","type2str":null,"dataname":"9.83 Million Chinese Japanese Bilingual Corpus for NLP and LLM Training","datazy":[{"title":"Format","desc":"Format","content":"TXT"},{"title":"Data content","desc":"Data content","content":"Chinese-Japanese parallel corpus"},{"title":"Data size","desc":"Data size","content":"9.83 million pairs of Chinese-Japanese Parallel Corpus Data."},{"title":"Language","desc":"Language","content":"Chinese, Japanese"},{"title":"Applications","desc":"Applications","content":"machine translation"},{"title":"Accuracy rate","desc":"Accuracy rate","content":"90%"}],"datatag":"Chinese,Japanese,Sino-Japan,Parallel corpus","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[{"name":"/data/apps/damp/temp/ziptemp/APY200214001_demo1711015206921/APY200214001_demo/APY200214001.jpeg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY200214001_demo1711015206921/APY200214001_demo/APY200214001.jpeg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=UWIrRqUw8h3Pnd7JBAu5O%2Bi2CRk%3D","intro":"","size":0,"progress":100,"type":"jpg"}],"officialSummary":"This dataset contains 9.83 million Chinese-Japanese sentence pairs stored in TXT format. The corpus covers multiple domains, including general topics, information technology, news, patents, and other specialized fields. Each sentence pair has undergone data anonymization and quality assurance processes to ensure usability for AI model development. It can be used as a basic corpus for text data analysis in fields such as machine translation.","dataexampl":null,"datakeyword":["chinese japanese parallel corpus","chinese japanese translation dataset","chinese japanese bilingual dataset","parallel corpus dataset","machine translation dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"nlu","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"}]","productNameEn":"9,830,000 Groups - Chinese-Japanese Parallel Corpus Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
9.83 Million Chinese Japanese Bilingual Corpus for NLP and LLM Training
chinese japanese parallel corpus
chinese japanese translation dataset
chinese japanese bilingual dataset
parallel corpus dataset
machine translation dataset
This dataset contains 9.83 million Chinese-Japanese sentence pairs stored in TXT format. The corpus covers multiple domains, including general topics, information technology, news, patents, and other specialized fields. Each sentence pair has undergone data anonymization and quality assurance processes to ensure usability for AI model development. It can be used as a basic corpus for text data analysis in fields such as machine translation.
This is a paid dataset licensed for commercial use. Ready-made datasets are available for immediate integration into AI projects.
Specifications
Format
TXT
Data content
Chinese-Japanese parallel corpus
Data size
9.83 million pairs of Chinese-Japanese Parallel Corpus Data.