Supervised Multimodal Bitransformers for Classifying Images and Text
Douwe Kiela Suvrat Bhooshan Hamed Firooz Ethan Perez Affiliation: Facebook AI; Davide Testuggine
Abstract
Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is often accompanied by other modalities such as images. We introduce a simple yet effective baseline for multimodal BERT-like architectures, a supervised multimodal bitransformer that jointly finetunes unimodally pretrained text and image encoders by projecting image embeddings to text token space. We approach or match state-of-the-art accuracy on several text-heavy multimodal classification tasks, outperforming strong baselines, including on hard test sets specifically designed to measure multimodal performance. Surprisingly, our method is competitive with ViLBERT, a self-supervised multimodal “BERT for vision-and-language” approach, while being much simpler and more easily extendible.
原文 arXiv:1909.02950;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1909.02950v2