ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images

Van Nguyen, Quan; Tran, Dan Quang; Pham, Huy Quang; Nguyen, Thang Kien-Bao; Nguyen, Nghia Hieu; Van Nguyen, Kiet; Nguyen, Ngan Luu-Thuy

Computer Science > Computation and Language

arXiv:2404.10652 (cs)

[Submitted on 16 Apr 2024 (v1), last revised 16 May 2025 (this version, v3)]

Title:ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images

Authors:Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

View PDF HTML (experimental)

Abstract:Visual Question Answerinng (VQA) is a complicated task that requires the capability of simultaneously processing natural language and images. This task was initially researched with a focus on developing methods to help machines understand objects and scene contexts in images. However, some scene text that carries explicit information about the full content of the image is not mentioned. Along with the continuous development of the AI era, there have been many studies on the reading comprehension ability of VQA models in the world. Therefore, we introduce the first large-scale dataset in Vietnamese specializing in the ability to understand scene text, we call it ViTextVQA (\textbf{Vi}etnamese \textbf{Text}-based \textbf{V}isual \textbf{Q}uestion \textbf{A}nswering dataset) which contains \textbf{over 16,000} images and \textbf{over 50,000} questions with answers. To tackle this task efficiently, we propose ViTextBLIP-2, an novel multimodal feature fusion Method, which optimizes Vietnamese OCR-based VQA by integrating a frozen Vision Transformer, SwinTextSpotter OCR, and ViT5 LLM with a trainable Q-Former for multimodal feature fusion. Through experiments with various state-of-the-art models, we uncover the significance of the order in which tokens in OCR text are processed and selected to formulate answers. This finding helped us significantly improve the performance of the baseline models on the ViTextVQA dataset. Our dataset is available (this https URL) for research purposes.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2404.10652 [cs.CL]
	(or arXiv:2404.10652v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2404.10652

Submission history

From: Nghia Hieu Nguyen [view email]
[v1] Tue, 16 Apr 2024 15:28:30 UTC (30,193 KB)
[v2] Sun, 9 Feb 2025 09:22:55 UTC (30,193 KB)
[v3] Fri, 16 May 2025 16:56:46 UTC (30,702 KB)

Computer Science > Computation and Language

Title:ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators