[{"@type":"PropertyValue","name":"Data content","value":"202,735 images of slide(RGB, the content is clear), with content description and QAs in annotation document file"},{"@type":"PropertyValue","name":"Diversity","value":"the slide images have four types as structure chart, graph, flow chart and figure"},{"@type":"PropertyValue","name":"Label content","value":"description and QAs of content in slide"},{"@type":"PropertyValue","name":"File format","value":"the image data is JPG/PNG, and the annotation document format is Markdown"},{"@type":"PropertyValue","name":"Language","value":"the main text of slide image is Chinese or English, and the annotation document file is labeled according to the language of PPT image"}]
{"id":1525,"datatype":"1","titleimg":"https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp","type1":"226","type1str":null,"type2":"254","type2str":null,"dataname":"202K Bilingual Slide VQA Dataset for Document AI","datazy":[{"isCheckLength":true,"title":"Data content","content":"202,735 images of slide(RGB, the content is clear), with content description and QAs in annotation document file"},{"isCheckLength":true,"title":"Diversity","content":"the slide images have four types as structure chart, graph, flow chart and figure"},{"isCheckLength":true,"title":"Label content","content":"description and QAs of content in slide"},{"isCheckLength":true,"title":"File format","content":"the image data is JPG/PNG, and the annotation document format is Markdown"},{"isCheckLength":true,"title":"Language","content":"the main text of slide image is Chinese or English, and the annotation document file is labeled according to the language of PPT image"}],"datatag":"Slide,VQA,Document analyze","technologydoc":null,"downurl":null,"datainfo":null,"standard":null,"dataylurl":null,"flag":null,"publishtime":null,"createby":null,"createtime":null,"ext1":null,"samplestoreloc":null,"hosturl":null,"datasize":null,"industryPlan":null,"keyInformation":"","samplePresentation":[{"name":"200001_chn_middle.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810182037/200001_chn_middle.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=1HJmY52wVacPDaz0KqEvUQDia9k%3D","intro":"","size":840352,"progress":100,"type":"jpg"},{"name":"229019_chn_complex.png","url":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810182037/229019_chn_complex.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=r4hJT8qHoPBJU0UbCpE6dr4vorQ%3D","intro":"","size":749196,"progress":100,"type":"jpg"}],"officialSummary":"2This dataset contains 202,735 presentation slide images. The dataset covers four major visual types commonly found in presentation slides: structural diagrams, graphs, flowcharts, and figures. The main text of the slides is available in either Chinese or English. Each annotation file is labeled according to the language of the corresponding presentation slide. This dataset can be used in document intelligence, document understanding, visual question answering (VQA), and vision-language model training.","dataexampl":null,"datakeyword":["slide VQA dataset","slide VQA data","slide visual question answering dataset","presentation VQA dataset","slide understanding dataset"],"isDelete":null,"ids":null,"idsList":null,"datasetCode":null,"productStatus":null,"tagTypeEn":"Type","tagTypeZh":null,"website":null,"samplePresentationList":null,"datazyList":null,"keyInformationList":null,"dataexamplList":null,"bgimg":null,"datazyScriptList":null,"datakeywordListString":null,"sourceShowPage":"llm","dataShowType":"[{\"code\":\"0\",\"language\":\"ZH\"},{\"code\":\"1\",\"language\":\"ZH\"},{\"code\":\"2\",\"language\":\"EN,PT,DE,KO,FR,ES,JP\"},{\"code\":\"3\",\"language\":\"EN\"},{\"code\":\"4\",\"language\":\"JP\"}]","productNameEn":"202,735 Images - Slide VQA Dataset","BGimg":"","voiceBg":["/shujutang/static/image/comm/audio_bg.webp","/shujutang/static/image/comm/audio_bg2.webp","/shujutang/static/image/comm/audio_bg3.webp","/shujutang/static/image/comm/audio_bg4.webp","/shujutang/static/image/comm/audio_bg5.webp"]}
https://www.nexdata.ai/shujutang/static/image/index/datatang_tuxiang_default.webp
[{"@type":"ImageObject","embedUrl":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810182037/200001_chn_middle.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=1HJmY52wVacPDaz0KqEvUQDia9k%3D"},{"@type":"ImageObject","embedUrl":"https://storage-product.datatang.com/damp/product/sample_presentation/20260810182037/229019_chn_complex.png?Expires=4102415999&OSSAccessKeyId=LTAI5tEBeSWUJiqjXvBMsxEu&Signature=r4hJT8qHoPBJU0UbCpE6dr4vorQ%3D"}]
202K Bilingual Slide VQA Dataset for Document AI
slide VQA dataset
slide VQA data
slide visual question answering dataset
presentation VQA dataset
slide understanding dataset
2This dataset contains 202,735 presentation slide images. The dataset covers four major visual types commonly found in presentation slides: structural diagrams, graphs, flowcharts, and figures. The main text of the slides is available in either Chinese or English. Each annotation file is labeled according to the language of the corresponding presentation slide. This dataset can be used in document intelligence, document understanding, visual question answering (VQA), and vision-language model training.
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
![Specifications]()
Specifications
Data content
202,735 images of slide(RGB, the content is clear), with content description and QAs in annotation document file
Diversity
the slide images have four types as structure chart, graph, flow chart and figure
Label content
description and QAs of content in slide
File format
the image data is JPG/PNG, and the annotation document format is Markdown
Language
the main text of slide image is Chinese or English, and the annotation document file is labeled according to the language of PPT image
![Sample]()
Sample
Tell Us Your Special Needs

What can Nexdata’s LLM datasets be used for?

Nexdata’s LLM datasets can support a wide range of large language model development tasks, including pre-training, supervised fine-tuning, instruction tuning, preference optimization, evaluation, and domain-specific model development. Depending on the dataset, data may include text, instruction-response pairs, conversations, question-answer pairs, reasoning data, and other structured or annotated content.

Can Nexdata customize LLM datasets based on our specific model and requirements?

Yes. If our off-the-shelf LLM datasets do not fully match your requirements, Nexdata provides flexible custom data collection, generation, annotation, curation, and quality control services. We can customize datasets based on your target languages, domains, use cases, data formats, task types, volume, and quality requirements to support specific LLM training and evaluation projects.

Can Nexdata provide large-scale and high-quality data for LLM development?

Yes. Nexdata can support large-scale LLM data projects across multiple languages, domains, and data types. Our data services include multi-stage quality control, data cleaning, annotation, validation, and curation to help ensure consistency and usability. For projects requiring data beyond our existing datasets, our customized data services can be scaled according to the required volume, specifications, and delivery schedule.
e264bb86-b898-483a-b820-c89adbc54e3d