Yin and Yang: Balancing and Answering Binary Visual Questions
Peng Zhang Thanks: The first two authors contributed equally. Affiliation: Virginia Tech Affiliation: {zhangp, ygoyal, dbatra, Yash Goyal Affiliation: Virginia Tech Affiliation: {zhangp, ygoyal, dbatra, Douglas Summers-Stay Affiliation: Army Research Laboratory Affiliation: Dhruv Batra Affiliation: Virginia Tech Affiliation: {zhangp, ygoyal, dbatra, Devi Parikh Affiliation: Virginia Tech Affiliation: {zhangp, ygoyal, dbatra,
Abstract
The complex compositional structure of language makes problems at the intersection of vision and language challenging. But language also provides a strong prior that can result in good superficial performance, without the underlying models truly understanding the visual content. This can hinder progress in pushing state of art in the computer vision aspects of multi-modal AI.
原文 arXiv:1511.05099;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1511.05099v5