en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

1,586,458 Sets-Document OCR&Phrasing Data

OCR
Document
Structured parsing

1,586,458 Sets Document OCR and Structured Analysis Data, Including Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, etc . The annotated files include OCR annotations and structured analysis.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Data Size
1,586,458 sets, including 1466168 sets of basic analysis data, 15289 sets of precise standard analysis data, 5001 sets of precise analysis data, and 100000 sets of original manual documents.
Data Types
Analysis data (Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, Chinese Journals), chinese original manual documents
Data Format
The original document file format is PDF, the document image file format is. png, the OCR annotation file format is JSON, and the structured parsing file format is markdown(Tables and formulas are in Latex format or screenshot links)
Sample Sample
  • 1,586,458 Sets-Document OCR&Phrasing Data
  • 1,586,458 Sets-Document OCR&Phrasing Data
Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

429993b7-65c8-4271-a46d-4ddbf8da0ee0

a85e338e-8e4a-4b28-b80c-48a925e8a051