en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

20,011 Multilingual Scene Text Image Caption Dataset for OCR & Vision AI

scene text dataset
OCR dataset
OCR image dataset
image caption dataset
multilingual OCR dataset
vision language dataset
image-text dataset
VLM dataset

This dataset contains 20,011 multilingual image-text pairs featuring natural scene text across 14 Asian and European languages. The images were collected from real-world environments, including shop signs, road signs, posters, public signage, and other outdoor and commercial scenes, with diverse camera viewpoints. All image captions are provided in English and describe the text layout, textual content, color, and other visual attributes. The dataset is ideal for Optical Character Recognition (OCR), scene text recognition, Vision-Language Models (VLMs), multimodal large language models (MLLMs), image captioning, and visual understanding applications.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Data size
20,011 pictures, 20,011descriptions
Language distribution
Asian languages: Korean, Indonesian, Malay, Vietnamese, Thai, Chinese, Japanese European languages: French, German, Italian, Portuguese, Russian, Spanish, English
Collection environment
including store plaques, stop signs, posters, road signs, prompts and other scenes
Collection diversity
including 14 languages, various natural scenes, and multiple shooting angles
Data format
image format is .jpg, text format is .txt
Collection equipment
mobile phone, camera
Description language
English
Text length
in principle, 30~60 words, usually 3-5 sentences
Main description content
text arrangement, text content, color, scene
Main deAccuracy ratescription content
the proportion of correctly labeled images is not less than 97%
Sample Sample
  • 20,011 Multilingual Scene Text Image Caption Dataset for OCR & Vision AI
  • 20,011 Multilingual Scene Text Image Caption Dataset for OCR & Vision AI
  • 20,011 Multilingual Scene Text Image Caption Dataset for OCR & Vision AI
Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

a522e04c-7e98-4190-ba1a-7314f8bc4a9a

f1e20527-4cde-436f-9ead-186dc857895e