[{"@type":"PropertyValue","name":"Data size","value":"128,900 images, one annotation document per image, 2,570,135 bounding boxes totally"},{"@type":"PropertyValue","name":"Data formats","value":"the format of image data includes .jpg and .jpeg, the format of annotation document is .json"},{"@type":"PropertyValue","name":"Scenes","value":"blur, nature"},{"@type":"PropertyValue","name":"Special shooting styles","value":"handheld shooting, moire"},{"@type":"PropertyValue","name":"Languages","value":"Arabic, French, German, Indian, Italian, Portuguese, Spanish"},{"@type":"PropertyValue","name":"Annotation","value":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}]
{"id":1416,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"128,900 Images-Multiple Scenes OCR Data of Seven Languages","datazy":[{"title":"Data size","content":"128,900 images, one annotation document per image, 2,570,135 bounding boxes totally"},{"title":"Data formats","content":"the format of image data includes .jpg and .jpeg, the format of annotation document is .json"},{"title":"Scenes","content":"blur, nature"},{"title":"Special shooting styles","content":"handheld shooting, moire"},{"title":"Languages","content":"Arabic, French, German, Indian, Italian, Portuguese, Spanish"},{"title":"Annotation","content":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}],"datatag":"Seven Languages,Multiple scenes,OCR data","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ara-blu-00001_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ara-blu-00001_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=vOyX12zcgd7jbSKYcoEYre7sqlE%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Fre-sho-00011_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Fre-sho-00011_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=BwUiilwEFYTCirEeSBrv80qgPFM%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ind-mul-011808_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ind-mul-011808_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=arws3uPhSUgLKFZIzEbe1fnvMpo%3D","intro":"","size":0,"progress":100,"type":"jpg"}],"officialSummary":"128,900 Images-Multiple scenes OCR Data of Seven Languages. This data have seven languages, including Arabic, French, German, Hindi, Italian, Portuguese, and Spanish. There are two scenes of this data like blur and nature, and two special shooting styles like handheld shooting and moire. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for OCR tasks.","dataexampl":null,"datakeyword":["Seven Languages","Multiple scenes","OCR data"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"128,900 Images-Multiple Scenes OCR Data of Seven Languages","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"],"firstList":[{"name":"/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ita-moi-00225_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102001_demo1706176800226/demo/Ita-moi-00225_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=qPE6WHz30aLe3K6jAtEav7%2B%2BXSc%3D","intro":"","size":0,"progress":100,"type":"jpg"}]}
128,900 Images-Multiple Scenes OCR Data of Seven Languages
Seven Languages
Multiple scenes
OCR data
128,900 Images-Multiple scenes OCR Data of Seven Languages. This data have seven languages, including Arabic, French, German, Hindi, Italian, Portuguese, and Spanish. There are two scenes of this data like blur and nature, and two special shooting styles like handheld shooting and moire. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for OCR tasks.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
128,900 images, one annotation document per image, 2,570,135 bounding boxes totally
Data formats
the format of image data includes .jpg and .jpeg, the format of annotation document is .json