Retrieval-Augmented Multimodal Language Modeling
Michihiro Yasunaga, Affiliation: Stanford University Correspondence to: Armen Aghajanyan, Affiliation: Meta AI Weijia Shi, Affiliation: University of Washington Rich James, Affiliation: Meta AI Jure Leskovec, Affiliation: Stanford University Percy Liang Affiliation: Stanford University Mike Lewis, Affiliation: Meta AI Luke Zettlemoyer, Affiliation: Meta AI Affiliation: University of Washington Wen-tau Yih Affiliation: Meta AI
Abstract
Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all their knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). Specifically, for the retriever, we use a pretrained CLIP, and for the generator, we train a CM3 Transformer on the LAION dataset. Our resulting model, named Retrieval-Augmented CM3 (RA-CM3), is the first multimodal model that can retrieve and generate both text and images. We show that RA-CM3 significantly outperforms baseline multimodal models such as DALL-E and CM3 on both image and caption generation tasks (12 FID and 17 CIDEr improvements on MS-COCO), while requiring much less compute for training ( $<\!30\%$ of DALL-E). Moreover, we show that RA-CM3 exhibits novel capabilities, such as fai
原文 arXiv:2211.12561;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2211.12561v2