Otter: A Multi-Modal Model with In-Context Instruction Tuning
Joshua Adrian Cahyono Jingkang Yang Chunyuan Li Ziwei Liu 🖂 Thanks: $ˆ∗$Equal Contribution. Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang and Ziwei Liu are with the S-Lab, Nanyang Technological University. E-mail: {libo0013, yuanhan002, liangyu.chen, jinghao003, fpu001, jo0001no, jingkang001, Chunyuan Li is with the Microsoft Research, Redmond. E-mail:
Abstract
Recent advances in Large Multimodal Models (LMMs) have unveiled great potential as visual assistants. However, most existing works focus on responding to individual instructions or using previous dialogues for contextual understanding. There is little discussion on employing both images and text as in-context examples to enhance the instruction following capability. To bridge this gap, we introduce the Otter model to leverage both textual and visual in-context examples for instruction tuning. Specifically, Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant. Otter seamlessly processes multi-modal inputs, supporting modalities including text, multiple images, and dynamic video content. To support the training of Otter, we present the MIMIC-IT (MultI-Modal In-Context Instruction Tuning) dataset, which encompasses over 3 million multi-modal instruction-response pairs, including approximately 2.2 million unique instructions across a broad spectrum of images and videos. MIMIC-IT has been carefully curated to feature a diverse array of in-context examples for each entry. Comprehensive evaluations suggest that in
原文 arXiv:2305.03726;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2305.03726v2