Enhanced Visual Question Answering: A Comparative Analysis and Textual Feature Extraction Via Convolutions

Zhang, Zhilin

Computer Science > Computer Vision and Pattern Recognition

arXiv:2405.00479 (cs)

[Submitted on 1 May 2024]

Title:Enhanced Visual Question Answering: A Comparative Analysis and Textual Feature Extraction Via Convolutions

Authors:Zhilin Zhang

View PDF HTML (experimental)

Abstract:Visual Question Answering (VQA) has emerged as a highly engaging field in recent years, attracting increasing research efforts aiming to enhance VQA accuracy through the deployment of advanced models such as Transformers. Despite this growing interest, there has been limited exploration into the comparative analysis and impact of textual modalities within VQA, particularly in terms of model complexity and its effect on performance. In this work, we conduct a comprehensive comparison between complex textual models that leverage long dependency mechanisms and simpler models focusing on local textual features within a well-established VQA framework. Our findings reveal that employing complex textual encoders is not invariably the optimal approach for the VQA-v2 dataset. Motivated by this insight, we introduce an improved model, ConvGRU, which incorporates convolutional layers to enhance the representation of question text. Tested on the VQA-v2 dataset, ConvGRU achieves better performance without substantially increasing parameter complexity.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2405.00479 [cs.CV]
	(or arXiv:2405.00479v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2405.00479

Submission history

From: Zhilin Zhang [view email]
[v1] Wed, 1 May 2024 12:39:35 UTC (602 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Enhanced Visual Question Answering: A Comparative Analysis and Textual Feature Extraction Via Convolutions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Enhanced Visual Question Answering: A Comparative Analysis and Textual Feature Extraction Via Convolutions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators