en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

8K Arabic OCR Dataset for OCR Detection and Recognition Training

arabic OCR dataset
arabic text recognition dataset
scene text dataset
OCR training data
text detection dataset

This dataset contains 8,604 images collected from diverse real-world Arabic text scenes, covering various environments, shooting angles, and natural conditions. Each text instance is annotated with precise quadrilateral or polygon bounding boxes along with transcription labels, enabling accurate text localization and recognition. This data can be used for Arabic OCR applications and multilingual computer vision model training.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Data size
8,604 images, 65,231 Arabic quadrilateral bounding boxes, 909 Arabic polygon bounding boxes
Collecting environment
including shop plaque, stop board, poster, road sign, comic, prompt/reminder, warning, packing instruction, menu, building sign, magazine book covers, etc.
Data diversity
including a variety of natural scenes, multiple shooting angles
Device
cellphone, camera
Photographic angle
looking up angle, looking down angle, eye-level angle
Data format
the image data format is .jpg, the annotation file format is .json
Annotation content
line-level quadrilateral bounding box annotation and transcription for the texts, polygon bounding box annotation and transcription for the texts
Accuracy
the error bound of each vertex of quadrilateral or polygon bounding box is within 5 pixels, which is a qualified annotation, the accuracy of bounding boxes is not less than 95%; the texts transcription accuracy is not less than 95%
Sample Sample
  • 8K Arabic OCR Dataset for OCR Detection and Recognition Training
  • 8K Arabic OCR Dataset for OCR Detection and Recognition Training
  • 8K Arabic OCR Dataset for OCR Detection and Recognition Training
Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

8b40b6fb-177a-4738-a3c9-fb1dbf48f262

436f5212-2f75-46ba-886b-6c183bf17a5d