ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, Stefan Lee
Abstract
We thank the reviewers for the thoughtful feedback! We are encouraged that all voted to accept, finding the paper clear / well-organized [R1]; our approach “very interesting” [R3] and novel [R2 R3]; our results significant and well-demonstrated [R1 R2]; and likely to be built on by the community [R1]. We are pleased they recognized the value of transferring visio-linguistic pretraining [R1 R2 R3] and the demonstrated benefits of our co-attentional two-stream model over a direct extension of BERT [R2 R3]. We respond to select comments below but will address all feedback.
原文 arXiv:1908.02265;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1908.02265v1