en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

OCR Datasets

Nexdata provides high-quality OCR datasets for document recognition, scene text detection, handwriting recognition, and multilingual text extraction applications.

Data Type

All
37
Document
9
General Scenario
13
Handwriting
17
Internet image
2
Invoice
3
Others
5
Test paper
1
Table
1

Language

All
37
Chinese
9
English
10
Hindi
5
Japanese
10
Korean
8
Others
21
Vietnamese
4

Handwriting OCR Dataset – Japanese and Korean (22,163 Images)

This dataset contains handwritten text images collected from 100 individuals, including 50 Japanese, 49 Koreans and 1 Afghan. For different subjects, the corpus are different. The data diversity includes multiple cellphone models and different corpus. This dataset can be used for tasks such as handwriting OCR models, handwritten text recognition systems, and multilingual OCR pipelines
handwriting OCR dataset handwritten OCR dataset handwriting recognition dataset Korean handwriting OCR dataset multilingual handwriting OCR dataset Japanese handwriting OCR dataset

1.58M Document OCR Dataset with Structured Analysis

This dataset contains 1,586,458 document data sets covering a wide range of Chinese-language materials, including Chinese textbooks, e-books, and teaching reference materials. The annotation files provide both OCR annotations and structured document analysis. The dataset is suitable for document OCR, document understanding tasks.
document OCR dataset document image OCR dataset document understanding dataset document layout analysis dataset document structure dataset

9,574 Images – Multilingual Handwriting OCR Dataset (8 Languages)

This dataset includes 9,574 handwriting images across 8 languages, including English, Spanish Portuguese and more. The data diversity includes multiple collecting scenes, different text carriers and different photographic angles(looking up, eye-level, looking down). In terms of annotation, each text line is annotated with quadrilateral polygons and transcription. The dataset can be used for training and evaluating OCR models, handwriting recognition systems, and multilingual text extraction tasks in AI and computer vision.
handwriting OCR dataset handwritten text recognition data multi-language handwriting OCR data OCR training data polygon-annotated handwriting dataset

426,687 Images - Multilingual OCR Dataset – Document & Scene Text

426,687 high-resolution images featuring multilingual Optical Character Recognition (OCR) data across both natural scenes and various document types. This dataset spans 20 languages, including Traditional Chinese, Simplified Chinese, Japanese, Korean, Thai, Vietnamese, Indonesian, Malay, Polish, and more. The data covers a wide range of real-world conditions—natural scenes, printed documents, handwritten notes, signs, and posters—captured from multiple countries and environments, with varied backgrounds, lighting conditions, and camera angles. All images are annotated for OCR tasks, making this dataset highly suitable for training deep learning models for text detection, recognition, and layout analysis in multi-language scenarios. The dataset complies with global data protection standards (GDPR, CCPA, PIPL), and is validated by leading AI enterprises for commercial and research applications.
multilingual OCR dataset scene text dataset document OCR images Chinese OCR Japanese OCR dataset Thai text recognition OCR dataset with annotation multi-language OCR document image dataset OCR training data

5,162 Images – Traditional Chinese Handwriting OCR Dataset

This dataset contains 5,162 handwriting images from 262 individuals, covering Traditional Chinese characters used in Taiwan. Each text in the data were annotated with quadrilateral bounding boxes. The handwriting ocr data can be used for training and evaluating OCR models, Traditional Chinese character recognition systems, and AI-based handwriting applications. The accuracy of line-level annotation and transcription is >= 97%.
Traditional Chinese handwriting OCR dataset handwriting OCR dataset for Traditional Chinese Traditional Chinese handwriting recognition

5,000 Images-4 Categories Cards OCR Data

5,000 Images-4 Categories Cards OCR Data, applicable for tasks such as character recognition.
Multiple card types;OCR

102,870 Sets - Chinese Instruction Manual OCR&Parsing Data

102,870 Sets - Chinese Instruction Manual OCR&Parsing Data, including Drug instructions, Installation instructions, Operating Instructions, Product manual, User manual, the annotation files contain OCR annotations and structured parsing.
Instruction document PDF

14,980 Images PPT OCR Data of 8 Languages

14,980 Images PPT OCR Data of 8 Languages. This dataset includes 8 languages, multiple scenes, different photographic angles, different photographic distances, different light conditions. For annotation, line-level quadrilateral bounding box annotation and transcription for the texts were annotated in the data. The dataset can be used for tasks such as OCR of multi-language.
Multiple scenes Multiple languages Different photographic angles Different photographic distances Different light conditions

15,148 Sets - Japanese and Korean Menu OCR Data

15,148 Sets - Japanese and Korean Menu OCR Data, including 9,701 Japanese menus and 5,447 Korean menus, annotation content includes row level quadrilateral box annotation, row level content transcription (a small portion of data is column level quadrilateral box annotation, column level content transcription)
Menu Images Japanese Korean

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
191fde3b-f953-4e20-aca2-b5e83dc7154d