[{"@type":"PropertyValue","name":"Data size","value":"30,276 images, one annotation document per image, 500,579 bounding boxes totally"},{"@type":"PropertyValue","name":"Data formats","value":"the format of image data is .jpg, the format of annotation document is .json"},{"@type":"PropertyValue","name":"Text mediums","value":"A4 paper, lined paper"},{"@type":"PropertyValue","name":"Writing style","value":"horizontal left-to-right writing, including different handwriting styles, different text colors (black, blue, red)"},{"@type":"PropertyValue","name":"Language","value":"English"},{"@type":"PropertyValue","name":"Annotation","value":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}]
{"id":1417,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"30K English Handwriting OCR Dataset","datazy":[{"title":"Data size","content":"30,276 images, one annotation document per image, 500,579 bounding boxes totally"},{"title":"Data formats","content":"the format of image data is .jpg, the format of annotation document is .json"},{"title":"Text mediums","content":"A4 paper, lined paper"},{"title":"Writing style","content":"horizontal left-to-right writing, including different handwriting styles, different text colors (black, blue, red)"},{"title":"Language","content":"English"},{"title":"Annotation","content":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}],"datatag":"English,Handwriting,OCR data","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-50284_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-50284_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=P6rnWf2QAfUy37NdgZmxqnz7hno%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-00001_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-00001_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=3lTrzGDcgSfXW3L5XQYWR69Vyo8%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-14570_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102002_demo1712656803975/demo/Eng-han-14570_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=dKkTqS1G2%2FVkkH3M5Rz%2BQdXOHII%3D","intro":"","size":0,"progress":100,"type":"jpg"}],"officialSummary":"This dataset contains 30,276 images of English handwritten text, written horizontally from left to right. It covers diverse handwriting styles and text colors, including black, blue, and red. The handwriting is captured on two types of writing surfaces: A4 paper and lined paper. Text regions are annotated with polygonal bounding boxes for text localization, along with corresponding text transcriptions. This data can be used for English handwriting OCR tasks.","dataexampl":null,"datakeyword":["english handwriting OCR dataset","english handwriting dataset","handwriting recognition dataset","HTR dataset","handwriting OCR dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,KO\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"30,276 Images-English Handwriting OCR Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
This dataset contains 30,276 images of English handwritten text, written horizontally from left to right. It covers diverse handwriting styles and text colors, including black, blue, and red. The handwriting is captured on two types of writing surfaces: A4 paper and lined paper. Text regions are annotated with polygonal bounding boxes for text localization, along with corresponding text transcriptions. This data can be used for English handwriting OCR tasks.
This is a paid dataset licensed for commercial use. Ready-made datasets are available for immediate integration into AI projects.
Specifications
Data size
30,276 images, one annotation document per image, 500,579 bounding boxes totally
Data formats
the format of image data is .jpg, the format of annotation document is .json
Text mediums
A4 paper, lined paper
Writing style
horizontal left-to-right writing, including different handwriting styles, different text colors (black, blue, red)
Language
English
Annotation
polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription
Nexdata provides OCR datasets covering a wide range of data types, including documents, handwriting, invoices, test papers, and forms. The datasets cover diverse languages, layouts, text styles, and real-world scenarios to support OCR model training, text recognition, document analysis, and other document AI applications.
Can Nexdata customize OCR datasets for specific requirements?
Yes. If our off-the-shelf OCR datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and quality control services. We can customize datasets based on target languages, document types, handwriting styles, layouts, image conditions, data volume, and annotation specifications for specific OCR applications.
How does Nexdata ensure the quality and scalability of OCR datasets?
Nexdata applies multi-stage quality control throughout data collection, annotation, validation, and delivery. With extensive resources covering documents, handwriting, invoices, test papers, and forms, we can support both large-scale OCR projects and specialized datasets with customized formats, annotation requirements, and quality standards.