[{"@type":"PropertyValue","name":"Data Size","value":"1,586,458 sets, including 1466168 sets of basic analysis data, 15289 sets of precise standard analysis data, 5001 sets of precise analysis data, and 100000 sets of original manual documents."},{"@type":"PropertyValue","name":"Data Types","value":"Analysis data (Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, Chinese Journals), chinese original manual documents"},{"@type":"PropertyValue","name":"Data Format","value":"The original document file format is PDF, the document image file format is. png, the OCR annotation file format is JSON, and the structured parsing file format is markdown(Tables and formulas are in Latex format or screenshot links)"}]
{"id":1749,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"1,586,458 Sets-Document OCR&Phrasing Data","datazy":[{"title":"Data Size","content":"1,586,458 sets, including 1466168 sets of basic analysis data, 15289 sets of precise standard analysis data, 5001 sets of precise analysis data, and 100000 sets of original manual documents."},{"title":"Data Types","content":"Analysis data (Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, Chinese Journals), chinese original manual documents"},{"title":"Data Format","content":"The original document file format is PDF, the document image file format is. png, the OCR annotation file format is JSON, and the structured parsing file format is markdown(Tables and formulas are in Latex format or screenshot links)"}],"datatag":"OCR,Document,Structured parsing","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"page3_demo.png","url":"https://storage-product.datatang.com/damp/product/samplePresentation_ipad/20250314162603/page3_demo.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=cYTecq3HKUv5MkOgaXcuxKIl44g%3D","intro":"","size":3545030,"progress":100,"type":"jpg"},{"name":"page33_demo.png","url":"https://storage-product.datatang.com/damp/product/samplePresentation_ipad/20250314162603/page33_demo.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=zgsnZaJ7L6PXUF%2BtpsBVwuzRajM%3D","intro":"","size":1837171,"progress":100,"type":"jpg"},{"name":"000003.pdf","url":"https://storage-product.datatang.com/damp/product/samplePresentation_ipad/20250616114549/000003.pdf?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=iIVxZU5Vd9cEkPdo0jUDIHT%2FRwo%3D","intro":"原始PDF文件","size":84223,"progress":100,"type":"mp4"}],"officialSummary":"1,586,458 Sets Document OCR and Structured Analysis Data, Including Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, etc . The annotated files include OCR annotations and structured analysis.","dataexampl":null,"datakeyword":["OCR","Document","Structured parsing"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"1,586,458 Sets-Document OCR&Parsing Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"],"firstList":[{"name":"000003.png","url":"https://storage-product.datatang.com/damp/product/samplePresentation_ipad/20250616114549/000003.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=Lo7zMoOVV3kaFw3C8zaXdtosoSs%3D","intro":"md文件截图","size":26592,"progress":100,"type":"jpg"},{"name":"000002.png","url":"https://storage-product.datatang.com/damp/product/samplePresentation_ipad/20260810141347/000002.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=ZONotD%2FMB1J%2Fht0PVODySm%2FB6M0%3D","intro":"","size":572226,"progress":100,"type":"jpg"}]}
1,586,458 Sets Document OCR and Structured Analysis Data, Including Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, etc . The annotated files include OCR annotations and structured analysis.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data Size
1,586,458 sets, including 1466168 sets of basic analysis data, 15289 sets of precise standard analysis data, 5001 sets of precise analysis data, and 100000 sets of original manual documents.
Data Types
Analysis data (Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, Chinese Journals), chinese original manual documents
Data Format
The original document file format is PDF, the document image file format is. png, the OCR annotation file format is JSON, and the structured parsing file format is markdown(Tables and formulas are in Latex format or screenshot links)