[{"@type":"PropertyValue","name":"Data size","value":"5,000 images"},{"@type":"PropertyValue","name":"Collection diversity","value":"multiple card types"},{"@type":"PropertyValue","name":"Data format","value":"the image data format is commonly used in formats such as . jpg, and the annotation document format is .json"},{"@type":"PropertyValue","name":"Annotation Contents","value":"Row level quadrilateral box annotation, row level content rewriting (a small amount of data is column level quadrilateral box annotation, column level content rewriting)"},{"@type":"PropertyValue","name":"Quality Requirements","value":"The accuracy of image label naming exceeds 98%;The correct detection is when the vertices of a quadrilateral box do not exceed 5 pixels, and the accuracy of the detection box is not less than 95%;The accuracy of text transcription shall not be less than 95%.;"},{"@type":"PropertyValue","name":"Data resolution","value":"1,280*720 or more"}]
{"id":1427,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"5,000 Images-4 Categories Cards OCR Data","datazy":[{"title":"Data size","content":"5,000 images"},{"title":"Collection diversity","content":"multiple card types"},{"title":"Data format","content":"the image data format is commonly used in formats such as . jpg, and the annotation document format is .json"},{"title":"Annotation Contents","content":"Row level quadrilateral box annotation, row level content rewriting (a small amount of data is column level quadrilateral box annotation, column level content rewriting)"},{"title":"Quality Requirements","content":"The accuracy of image label naming exceeds 98%;The correct detection is when the vertices of a quadrilateral box do not exceed 5 pixels, and the accuracy of the detection box is not less than 95%;The accuracy of text transcription shall not be less than 95%.;"},{"title":"Data resolution","content":"1,280*720 or more"}],"datatag":"Multiple card types;OCR","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[],"officialSummary":"5,000 Images-4 Categories Cards OCR Data, applicable for tasks such as character recognition.","dataexampl":null,"datakeyword":["Multiple card types;OCR"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP,KO\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"5,000 Images-4 Categories Cards OCR Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
5,000 Images-4 Categories Cards OCR Data, applicable for tasks such as character recognition.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
5,000 images
Collection diversity
multiple card types
Data format
the image data format is commonly used in formats such as . jpg, and the annotation document format is .json
Annotation Contents
Row level quadrilateral box annotation, row level content rewriting (a small amount of data is column level quadrilateral box annotation, column level content rewriting)
Quality Requirements
The accuracy of image label naming exceeds 98%;The correct detection is when the vertices of a quadrilateral box do not exceed 5 pixels, and the accuracy of the detection box is not less than 95%;The accuracy of text transcription shall not be less than 95%.;
Nexdata provides OCR datasets covering a wide range of data types, including documents, handwriting, invoices, test papers, and forms. The datasets cover diverse languages, layouts, text styles, and real-world scenarios to support OCR model training, text recognition, document analysis, and other document AI applications.
Can Nexdata customize OCR datasets for specific requirements?
Yes. If our off-the-shelf OCR datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and quality control services. We can customize datasets based on target languages, document types, handwriting styles, layouts, image conditions, data volume, and annotation specifications for specific OCR applications.
How does Nexdata ensure the quality and scalability of OCR datasets?
Nexdata applies multi-stage quality control throughout data collection, annotation, validation, and delivery. With extensive resources covering documents, handwriting, invoices, test papers, and forms, we can support both large-scale OCR projects and specialized datasets with customized formats, annotation requirements, and quality standards.