[{"@type":"PropertyValue","name":"Data size","value":"1,000 people, each subject collects 14 images"},{"@type":"PropertyValue","name":"Population distribution","value":"gender distribution: 516 males, 484 females; age distribution: 19 people under 18 years old, 937 people from 18 to 45 years old, 31 people from 46 to 60 years old, 13 people over 60 years old"},{"@type":"PropertyValue","name":"Writer","value":"Europeans who often write Spanish"},{"@type":"PropertyValue","name":"Collecting environment","value":"pure color background"},{"@type":"PropertyValue","name":"Device","value":"scanner"},{"@type":"PropertyValue","name":"Photographic angle","value":"eye-level angle"},{"@type":"PropertyValue","name":"Data format","value":"the image data format is .png"},{"@type":"PropertyValue","name":"Data content","value":"including address, company name and personal name, each image has 20 writing boxes"},{"@type":"PropertyValue","name":"Accuracy rate","value":"The collection content accuracy is not less than 97%"}]
{"id":1405,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"1,000 People - Spanish Handwriting OCR Dataset","datazy":[{"title":"Data size","desc":"Data size","content":"1,000 people, each subject collects 14 images"},{"title":"Population distribution","desc":"Population distribution","content":"gender distribution: 516 males, 484 females; age distribution: 19 people under 18 years old, 937 people from 18 to 45 years old, 31 people from 46 to 60 years old, 13 people over 60 years old"},{"title":"Writer","desc":"Writer","content":"Europeans who often write Spanish"},{"title":"Collecting environment","desc":"Collecting environment","content":"pure color background"},{"title":"Device","desc":"Device","content":"scanner"},{"title":"Photographic angle","desc":"Photographic angle","content":"eye-level angle"},{"title":"Data format","desc":"Data format","content":"the image data format is .png"},{"title":"Data content","desc":"Data content","content":"including address, company name and personal name, each image has 20 writing boxes"},{"title":"Accuracy rate","desc":"Accuracy rate","content":"The collection content accuracy is not less than 97%"}],"datatag":"Spanish,Handwriting,OCR","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[],"officialSummary":"The writers are Europeans who often write spanish. The device is scanner, the collection angle is eye-level angle. The dataset content includes address, company name, personal name.The dataset can be used for Spanish OCR models and handwritten text recognition systems.","dataexampl":null,"datakeyword":["OCR training dataset","Spanish handwriting ocr dataset","Spanish ocr dataset","Spanish HTR Dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"1,000 People - Spanish Handwriting OCR Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
The writers are Europeans who often write spanish. The device is scanner, the collection angle is eye-level angle. The dataset content includes address, company name, personal name.The dataset can be used for Spanish OCR models and handwritten text recognition systems.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
1,000 people, each subject collects 14 images
Population distribution
gender distribution: 516 males, 484 females; age distribution: 19 people under 18 years old, 937 people from 18 to 45 years old, 31 people from 46 to 60 years old, 13 people over 60 years old
Writer
Europeans who often write Spanish
Collecting environment
pure color background
Device
scanner
Photographic angle
eye-level angle
Data format
the image data format is .png
Data content
including address, company name and personal name, each image has 20 writing boxes
Accuracy rate
The collection content accuracy is not less than 97%
Nexdata provides OCR datasets covering a wide range of data types, including documents, handwriting, invoices, test papers, and forms. The datasets cover diverse languages, layouts, text styles, and real-world scenarios to support OCR model training, text recognition, document analysis, and other document AI applications.
Can Nexdata customize OCR datasets for specific requirements?
Yes. If our off-the-shelf OCR datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and quality control services. We can customize datasets based on target languages, document types, handwriting styles, layouts, image conditions, data volume, and annotation specifications for specific OCR applications.
How does Nexdata ensure the quality and scalability of OCR datasets?
Nexdata applies multi-stage quality control throughout data collection, annotation, validation, and delivery. With extensive resources covering documents, handwriting, invoices, test papers, and forms, we can support both large-scale OCR projects and specialized datasets with customized formats, annotation requirements, and quality standards.