From Recognition to Cognition: Visual Commonsense Reasoning
Rowan Zellers♠ Yonatan Bisk♠ Ali Farhadi♠♡ Yejin Choi♠♡ ♠Paul G. Allen School of Computer Science、Engineering, University of Washington ♡Allen Institute for Artificial Intelligence visualcommonsense.com
Abstract
Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people’s actions, goals, and mental states. While this task is easy for humans, it is tremendously difficult for today’s vision systems, requiring higher-order cognition and commonsense reasoning about the world. We formalize this task as Visual Commonsense Reasoning. Given a challenging question about an image, a machine must answer correctly and then provide a rationale justifying its answer.
原文 arXiv:1811.10830;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1811.10830v2