[{"@type":"PropertyValue","name":"Data size","value":"29,954 images, including 8,798 images in Khmer (Cambodian), 11,575 images in Lao, and 9,581 images in Burmese"},{"@type":"PropertyValue","name":"Collecting environment","value":"Natural scenes: shop signs, posters, warnings, road signs, food packages, billboards, street views, etc. Document Photograph: cards, receipts, newspapers, books (documents, newspapers, books, test pap"},{"@type":"PropertyValue","name":"Data diversity","value":"multiple languages, multiple collection types, multiple shooting angles"},{"@type":"PropertyValue","name":"Device","value":"cellphone, computer"},{"@type":"PropertyValue","name":"Data format","value":"the image format is a common one such as .png"},{"@type":"PropertyValue","name":"Accuracy rate","value":"according to the collection requirements, the collection accuracy is not less than 95%"}]
{"id":1758,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"29,954 Images - Southeast Asian OCR Dataset (Khmer, Lao, Burmese)","datazy":[{"title":"Data size","content":"29,954 images, including 8,798 images in Khmer (Cambodian), 11,575 images in Lao, and 9,581 images in Burmese","desc":"Data size"},{"title":"Collecting environment","content":"Natural scenes: shop signs, posters, warnings, road signs, food packages, billboards, street views, etc. Document Photograph: cards, receipts, newspapers, books (documents, newspapers, books, test pap","desc":"Collecting environment"},{"title":"Data diversity","content":"multiple languages, multiple collection types, multiple shooting angles","desc":"Data diversity"},{"title":"Device","content":"cellphone, computer","desc":"Device"},{"title":"Data format","content":"the image format is a common one such as .png","desc":"Data format"},{"title":"Accuracy rate","content":"according to the collection requirements, the collection accuracy is not less than 95%","desc":"Accuracy rate"}],"datatag":"OCR,Southeast Asian Languages,Natural Scenes,Document Photograph,Electronic Scenes","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[],"officialSummary":"29,954 Images - OCR Collection Data of Southeast Asian Languages, including Khmer (Cambodia), Lao and Burmese. The diversity of collection includes multiple languages, multiple collection types, multiple shooting angles. This set of data can be used for Southeast Asian language OCR tasks.","dataexampl":null,"datakeyword":["Southeast Asian OCR dataset","Khmer OCR data","Lao OCR dataset","Burmese OCR dataset","natural scene OCR","minority language text recognition","multilingual OCR dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,PT,DE,KO,FR,ES\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"29,954 Images - OCR Collection Data of Southeast Asian Languages","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
29,954 Images - Southeast Asian OCR Dataset (Khmer, Lao, Burmese)
Southeast Asian OCR dataset
Khmer OCR data
Lao OCR dataset
Burmese OCR dataset
natural scene OCR
minority language text recognition
multilingual OCR dataset
29,954 Images - OCR Collection Data of Southeast Asian Languages, including Khmer (Cambodia), Lao and Burmese. The diversity of collection includes multiple languages, multiple collection types, multiple shooting angles. This set of data can be used for Southeast Asian language OCR tasks.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
29,954 images, including 8,798 images in Khmer (Cambodian), 11,575 images in Lao, and 9,581 images in Burmese
Collecting environment
Natural scenes: shop signs, posters, warnings, road signs, food packages, billboards, street views, etc. Document Photograph: cards, receipts, newspapers, books (documents, newspapers, books, test pap
Nexdata provides OCR datasets covering a wide range of data types, including documents, handwriting, invoices, test papers, and forms. The datasets cover diverse languages, layouts, text styles, and real-world scenarios to support OCR model training, text recognition, document analysis, and other document AI applications.
Can Nexdata customize OCR datasets for specific requirements?
Yes. If our off-the-shelf OCR datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, annotation, and quality control services. We can customize datasets based on target languages, document types, handwriting styles, layouts, image conditions, data volume, and annotation specifications for specific OCR applications.
How does Nexdata ensure the quality and scalability of OCR datasets?
Nexdata applies multi-stage quality control throughout data collection, annotation, validation, and delivery. With extensive resources covering documents, handwriting, invoices, test papers, and forms, we can support both large-scale OCR projects and specialized datasets with customized formats, annotation requirements, and quality standards.