Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
Abhishek Das1, Satwik Kottur, José M.F. Moura2, Stefan Lee3, Dhruv Batra1 1Georgia Institute of Technology, 2Carnegie Mellon University, 3Virginia Tech visualdialog.org The first two authors (AD, SK) contributed equally.
Abstract
We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative ‘image guessing’ game between two agents – Q-bot and A-bot– who communicate in natural language dialog so that Q-bot can select an unseen image from a lineup of images. We use deep reinforcement learning (RL) to learn the policies of these agents end-to-end – from pixels to multi-agent multi-round dialog to game reward.
原文 arXiv:1703.06585;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1703.06585v2