[{"@type":"PropertyValue","name":"Data size","value":"19,101 images, including 9,992 images of subtitles, 9,109 images of news headlines"},{"@type":"PropertyValue","name":"Data diversity","value":"including multiple films and television series, 14 types of news program"},{"@type":"PropertyValue","name":"Language distribution","value":"Chinese, English"},{"@type":"PropertyValue","name":"Data format","value":"the image data format is .jpg, the annotation file format is .json"},{"@type":"PropertyValue","name":"Annotation content","value":"line-level rectangular bounding box annotation and transcription for the texts"},{"@type":"PropertyValue","name":"Accuracy","value":"the error bound of each vertex of a rectangular bounding box is within 5 pixels, which is a qualified annotation, the accuracy of bounding boxes is not less than 97%; the texts transcription accuracy is not less than 97%"}]
{"id":155,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"147","type1str":null,"type2":"150","type2str":null,"dataname":"19,101 Images – Subtitles and News Headlines OCR Data","datazy":[{"title":"Data size","content":"19,101 images, including 9,992 images of subtitles, 9,109 images of news headlines"},{"title":"Data diversity","content":"including multiple films and television series, 14 types of news program"},{"title":"Language distribution","content":"Chinese, English"},{"title":"Data format","content":"the image data format is .jpg, the annotation file format is .json"},{"title":"Annotation content","content":"line-level rectangular bounding box annotation and transcription for the texts"},{"title":"Accuracy","content":"the error bound of each vertex of a rectangular bounding box is within 5 pixels, which is a qualified annotation, the accuracy of bounding boxes is not less than 97%; the texts transcription accuracy is not less than 97%"}],"datatag":"Multiple films and television series,14 types of news program","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":null,"samplePresentation":[],"officialSummary":"19,101 Images – Subtitles and News Headlines OCR Data. The data includes 9,992 images of subtitles and 9,109 images of news headlines. The data diversity includes multiple films and television series, 14 types of news program. For annotation, line-level rectangular bounding box annotation and transcription for the texts were adopted for the subtitles in images and news headlines in images. The data can be used for OCR tasks of subtitles and news headlines.","dataexampl":null,"datakeyword":["Multiple films and television series","14 types of news program"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Data Type,Language","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"ocr","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,JP\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"19,101 Images – Subtitles and News Headlines OCR Data","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
19,101 Images – Subtitles and News Headlines OCR Data
Multiple films and television series
14 types of news program
19,101 Images – Subtitles and News Headlines OCR Data. The data includes 9,992 images of subtitles and 9,109 images of news headlines. The data diversity includes multiple films and television series, 14 types of news program. For annotation, line-level rectangular bounding box annotation and transcription for the texts were adopted for the subtitles in images and news headlines in images. The data can be used for OCR tasks of subtitles and news headlines.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
Specifications
Data size
19,101 images, including 9,992 images of subtitles, 9,109 images of news headlines
Data diversity
including multiple films and television series, 14 types of news program
Language distribution
Chinese, English
Data format
the image data format is .jpg, the annotation file format is .json
Annotation content
line-level rectangular bounding box annotation and transcription for the texts
Accuracy
the error bound of each vertex of a rectangular bounding box is within 5 pixels, which is a qualified annotation, the accuracy of bounding boxes is not less than 97%; the texts transcription accuracy is not less than 97%