en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

9,497 Images - OCR Data of 10 Types of Forms

OCR
forms

9,497 Images - OCR Data of 10 Types of Forms. Rectangular bounding boxes were adopted to annotate forms. The data can be used for tasks such as forms detection.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Date size
9,497 images, 10 types of forms
Collecting environment
pure color background
Data diversity
multiple types of forms
Data format
the image data format is .jpg , the annotation file format is .json
Annotation content
rectangular bounding boxes of forms
Accuracy
The error bound of each vertex of rectangular bounding box is within 5 pixels, which is a qualified annotation, the accuracy of rectangular bounding boxes is not less than 95%
Sample Sample
  • 9,497 Images - OCR Data of 10 Types of Forms
  • 9,497 Images - OCR Data of 10 Types of Forms
  • 9,497 Images - OCR Data of 10 Types of Forms
Recommended DatasetsRecommended Dataset
262 People - 5,162 Images Handwriting OCR Data of Traditional Chinese Characters (Taiwan, China)

262 People - 5,162 Images Handwriting OCR Data of Traditional Chinese Characters (Taiwan, China). Texts in the data were annotated for the line-level quadrilateral bounding box. The handwriting ocr data can be used for traditional Chinese characters recognition application.The accuracy of line-level annotation and transcription is >= 97%.

Handwriting OCR eye-level angle line-level quadrilateral bounding box annotation and transcription for the texts traditional Chinese characters (Taiwan China) chinese ocr handwriting ocr ocr chinese OCR Training dataset Optical character recognition TextOCR Dataset Reading everyday scenes standard Chinese character Handprint Applicationrn
101 People - 4,538 Images Japanese Handwriting OCR Data

101 People - 4,538 Images Japanese Handwriting OCR Data. The text carrier is A4 paper. The dataset content includes social livelihood, entertainment, tour, sport, movie, composition and other fields. For annotation, character-level rectangular bounding box annotation and text transcription and line-level rectangular bounding box annotation and text transcription were adopted. The dataset can be used for tasks such as Japanese handwriting OCR.

Japanese handwriting OCR A4 paper scanner character-level rectangular bounding box annotation text transcription writing hand script calligraphy penmanship longhand graphology scrawl scribble manuscript autography chirography autograph pencraft fist penscript scription writings cursive print hieroglyphics write lettering manuscription shorthand chicken scratch scratching handwritten printing copperplate griffonage handwrite letter scrivenery scrivening calligraph cuneiform scripts signature book cursive script handwritings inscription pen running hand steno undersignature notation autographing calligraphic'
105,941 Images Natural Scenes OCR Data of 12 Languages

105,941 Images Natural Scenes OCR Data of 12 Languages. The data covers 12 languages (6 Asian languages, 6 European languages), multiple natural scenes, multiple photographic angles. For annotation, line-level quadrilateral bounding box annotation and transcription for the texts were annotated in the data. The data can be used for tasks such as OCR of multi-language.

Japanese Korean Indonesian Malay Vietnamese Thai French German Italian Portuguese Russian Spanish OCR natural scenes multiple photographic angles line-level quadrilateral bounding box annotation and transcription for the texts
17,561 Images of Primary School Mathematics Papers

17,561 Images of Primary School Mathematics Papers Collection Data. The data background is pure color. The data covers multiple question types, multiple types of test papers (math workbooks, test papers, competition test papers, etc.) and multiple grades. The data can be used for tasks such as intelligent scoring and homework guidance for primary school students.

Primary School Mathematics Papers OCR multiple types of questions (Vertical calculation Horizontal calculation Recursive calculation Fraction Solving equation etc.) multiple types of test papers (math workbooks test papers competition test questions etc.) multiple grades
4,995 Vietnamese OCR Images Data - Images with Annotation and Transcription

4,995 Vietnamese OCR Images Data - Images with Annotation and Transcription. The data includes 258 images of natural scenes, 2,553 Internet images, 2,184 document images. For line-level content annotation, line-level quadrilateral bounding box annotation and test transcription was adpoted; for column-level content annotation, column-level quadrilateral bounding box annotation and text transcription was adpoted. The data can be used for tasks such as Vietnamese recognition in multiple scenes.

Vietnamese OCR document images Internet images natural scenes multiple angles different light conditions quadrilateral bounding box annotation line-level transcription for the texts column-level transcription for the texts
3,506 Hindi OCR Images Data - Images with Annotation and Transcription

3,506 Hindi OCR Images Data - Images with Annotation and Transcription. The data includes 2,056 images of natural scenes, 1,103 Internet images and 347 document images. For line-level content annotation, line-level quadrilateral bounding box annotation and test transcription was adpoted; for column-level content annotation, column-level quadrilateral bounding box annotation and text transcription was adpoted. The data can be used for tasks such as Hindi character recognition in multiple scenes.

Hindi OCR document images Internet images natural scenes multiple angles different light conditions quadrilateral bounding box annotation line-level transcription for the texts column-level transcription for the texts
4,601 Images-22 Kinds of Bills OCR Data

4,601 Images-22 Kinds of Bills OCR Data. The data background is pure color. The data covers 22 kinds of bills of multiple provinces. In terms of annotation, line-level quadrilateral bounding box annotation, line-level transcription for the texts were annotated in the data. The data can be used for tasks such as OCR for bills.

OCR bill annotation bill transcription multiple types of bills multiple provinces
14,980 Images PPT OCR Data of 8 Languages

14,980 Images PPT OCR Data of 8 Languages. This dataset includes 8 languages, multiple scenes, different photographic angles, different photographic distances, different light conditions. For annotation, line-level quadrilateral bounding box annotation and transcription for the texts were annotated in the data. The dataset can be used for tasks such as OCR of multi-language.

PPT OCR meeting room conference room different photographic angles different photographic distances different light conditions line-level quadrilateral bounding box annotation and transcription for the texts
Tell Us Your Special Needs

By submitting, I agree to the Privacy Protection

d2aa5b1d-aa7c-4606-8e62-e81e205f37ff

72e6a833-ddbf-4294-877b-4abc4ff9403b