en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

OCR Datasets

Nexdata provides high-quality OCR datasets for document recognition, scene text detection, handwriting recognition, and multilingual text extraction applications.

Data Type

All
31
Document
6
General Scenario
12
Handwriting
16
Internet image
1
Invoice
3
Others
5
Test paper
1
Table
1

Language

All
31
Chinese
6
English
6
Hindi
5
Japanese
7
Korean
6
Others
19
Vietnamese
4

Form OCR Dataset – 9,497 Images of 10 Form Types

This dataset contains 9,497 images of 10 types of forms. Rectangular bounding boxes were adopted to annotate forms. The dataset can be used for tasks such as forms detection.
form OCR dataset document form image dataset OCR forms dataset form recognition dataset forms detection dataset

128,900 Images-Multiple Scenes OCR Data of Seven Languages

128,900 Images-Multiple scenes OCR Data of Seven Languages. This data have seven languages, including Arabic, French, German, Hindi, Italian, Portuguese, and Spanish. There are two scenes of this data like blur and nature, and two special shooting styles like handheld shooting and moire. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for OCR tasks.
Seven Languages Multiple scenes OCR data

1,586,458 Sets-Document OCR&Phrasing Data

1,586,458 Sets Document OCR and Structured Analysis Data, Including Chinese Textbooks, Chinese E-books, Chinese Teaching Reference Books, etc . The annotated files include OCR annotations and structured analysis.
OCR Document Structured parsing

9,401 Images-English Document OCR Data

9,401 Images-English Document OCR Data. The language and text contents of this data are English and play script, book, exam paper, etc. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for English document OCR tasks.
English Document OCR data

30,276 Images-English Handwriting OCR Data

30,276 Images-English Handwriting OCR Data. The language and writing style of this data are English and horizontal left-to-right writing, including different handwriting styles, different text colors (black, blue, red). There are two text mediums of this data like A4 paper and lined paper. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for English handwriting OCR tasks.
English Handwriting OCR data

5,156 Images - Handwritten Mathematical Formula OCR Dataset

This dataset contains 5,156 images of handwritten mathematical formulas collected under diverse writing conditions, including A4 paper, grid paper, lined paper, whiteboards, and other writing surfaces. The dataset covers various mathematical expressions, handwriting styles, paper types, and capture conditions. Images were collected from multiple viewpoints, including upward-facing and eye-level camera angles. The dataset is designed for Math OCR, handwritten formula recognition, mathematical expression recognition, and AI model training for education and document intelligence applications.
math OCR dataset handwritten formula dataset mathematical formula recognition dataset formula recognition dataset handwritten math recognition dataset

8K Arabic OCR Dataset for OCR Detection and Recognition Training

This dataset contains 8,604 images collected from diverse real-world Arabic text scenes, covering various environments, shooting angles, and natural conditions. Each text instance is annotated with precise quadrilateral or polygon bounding boxes along with transcription labels, enabling accurate text localization and recognition. This data can be used for Arabic OCR applications and multilingual computer vision model training.
arabic OCR dataset arabic text recognition dataset scene text dataset OCR training data text detection dataset

Vertical Text OCR Dataset with 57,645 Images and Polygon Annotations

This dataset contains 57,645 images collected from diverse real-world text scenes, including street scenes, shop signs, billboards, posters, decorative text, art lettering, and magazine covers. The dataset includes Chinese text and a small amount of English text. Each image is annotated with text localization information and transcription labels, including polygon and quadrilateral bounding boxes for vertical and non-vertical text regions. This dataset is suitable for tasks such as scene text OCR tasks and multi-oriented text recognition.
ocr dataset ocr training data scene text dataset text recognition dataset text detection dataset

Vietnamese OCR Dataset with Annotations and Transcriptions (4,995 Images)

This dataset contains 4,995 Vietnamese OCR images with annotations and text transcriptions. The data includes 258 natural scene images, 2,553 Internet images, and 2,184 document images. For line-level content annotation, quadrilateral bounding box annotations and text transcriptions are provided. For column-level content annotation, column-level quadrilateral bounding box annotation and text transcription are provided. The data can be used for tasks such as Vietnamese recognition in multiple scenes.
Vietnamese OCR dataset Vietnamese text recognition dataset Vietnamese OCR images Vietnamese OCR training data Vietnamese text detection dataset

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
83055705-76c5-4b2f-8aa6-06eea5d44e63