Quantifying and Alleviating the Language Prior Problem in Visual Question AnsweringConference: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval; July 21–25, 2019; Paris, FranceProceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’19), July 21–25, 2019, Paris, FrancePrice: 15.00DOI: 10.1145/3331184.3331186ISBN: 978-1-4503-6172-9/19/07Thanks: Corresponding Author: Zhiyong Cheng and Liqiang Nie.CCS: Information systems Information retrievalCCS: Computing methodologies Natural language processingCCS: Computing methodologies Computer vision
Yangyang Guo†, Zhiyong Cheng§, Liqiang Nie†, Yibing Liu†, Yinglong Wang§, Mohan Kankanhalli‡ Affiliation: †School of Computer Science and Technology, Shandong University Affiliation: §Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences) Affiliation: ‡School of Computing, National University of Singapore email: guoyang.eric, jason.zy.cheng, nieliqiang,
Abstract
Benefiting from the advancement of computer vision, natural language processing and information retrieval techniques, visual question answering (VQA), which aims to answer questions about an image or a video, has received lots of attentions over the past few years. Although some progress has been achieved so far, several studies have pointed out that current VQA models are heavily affected by the language prior problem, which means they tend to answer questions based on the co-occurrence patterns of question keywords (e.g., how many) and answers (e.g., 2) instead of understanding images and questions. Existing methods attempt to solve this problem by either balancing the biased datasets or forcing models to better understand images. However, only marginal effects and even performance deterioration are observed for the first and second solution, respectively. In addition, another important issue is the lack of measurement to quantitatively measure the extent of the language prior effect, which severely hinders the advancement of related techniques.
原文 arXiv:1905.04877;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1905.04877v1