[{"@type":"PropertyValue","name":"Data size","value":"9,401 images, one annotation document per image, 269,725 bounding boxes totally"},{"@type":"PropertyValue","name":"Data formats","value":"the format of image data is .jpg, the format of annotation document is .json"},{"@type":"PropertyValue","name":"Text contents","value":"play script, book, exam paper"},{"@type":"PropertyValue","name":"Language","value":"English"},{"@type":"PropertyValue","name":"Annotation","value":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}]
{"id":1418,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"9,401 Images-English Document OCR Data","datazy":[{"title":"Data size","content":"9,401 images, one annotation document per image, 269,725 bounding boxes totally"},{"title":"Data formats","content":"the format of image data is .jpg, the format of annotation document is .json"},{"title":"Text contents","content":"play script, book, exam paper"},{"title":"Language","content":"English"},{"title":"Annotation","content":"polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription"}],"datatag":"English,Document,OCR data","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[{"name":"/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-04600_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-04600_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=DxEBPqjqNijzWOTuywXaKuJiUjY%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-00001_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-00001_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=iK44MVQPp1ZiXb%2BZmB0OvNA8oQA%3D","intro":"","size":0,"progress":100,"type":"jpg"},{"name":"/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-09780_demo.jpg","url":"https://bj-oss-datatang-03.oss-cn-beijing.aliyuncs.com/filesInfoUpload/data/apps/damp/temp/ziptemp/APY240102003_demo1706176802203/demo/Eng-pun-09780_demo.jpg?Expires=4102329599&OSSAccessKeyId=LTAI8NWs2pDolLNH&Signature=abYNumZGoYcTTd9ENV3AJdqOClg%3D","intro":"","size":0,"progress":100,"type":"jpg"}],"officialSummary":"9,401 Images-English Document OCR Data. The language and text contents of this data are English and play script, book, exam paper, etc. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for English document OCR tasks.","dataexampl":null,"datakeyword":["English","Document","OCR data"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"9,401 Images-English Document OCR Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
9,401 Images-English Document OCR Data. The language and text contents of this data are English and play script, book, exam paper, etc. The annotation of this data includes polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription. This data can be used for English document OCR tasks.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
9,401 images, one annotation document per image, 269,725 bounding boxes totally
Data formats
the format of image data is .jpg, the format of annotation document is .json
Text contents
play script, book, exam paper
Language
English
Annotation
polygonal bounding box labeling of text (with a precision between rectangular bounding box labeling and image segmentation labeling) and transcription