en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

9,401 Images-English Document OCR Data

English
Document
OCR data

9,401 Images-English Document OCR Data. The language and text contents of this data are English and play script, book, exam paper, etc. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for English document OCR tasks.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Data size
9,401 images, one annotation document per image, 269,725 bounding boxes totally
Data formats
the format of image data is .jpg, the format of annotation document is .json
Text contents
play script, book, exam paper
Language
English
Annotation
polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription
Sample Sample
  • 9,401 Images-English Document OCR Data
  • 9,401 Images-English Document OCR Data
  • 9,401 Images-English Document OCR Data
Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

457c92bc-37b1-4589-b605-fc365f0fec33

fa3c390d-6ce4-499b-8e99-c65035efbd5e