LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu Affiliation: Dept. of Comp. Sci.、Tech., Institute for AI, BNRist, Tsinghua University Hao Cheng Affiliation: Microsoft Research, Redmond Haotian Liu Hao Zhang Feng Li Tianhe Ren Affiliation: University of Wisconsin-Madison HKUST IDEA Research Xueyan Zou Jianwei Yang Affiliation: Microsoft Research, Redmond Hang Su Affiliation: Dept. of Comp. Sci.、Tech., Institute for AI, BNRist, Tsinghua University Jun Zhu Affiliation: Dept. of Comp. Sci.、Tech., Institute for AI, BNRist, Tsinghua University Lei Zhang Affiliation: University of Wisconsin-Madison HKUST IDEA Research Jianfeng Gao Affiliation: Microsoft Research, Redmond Chunyuan Li Affiliation: Microsoft Research, Redmond Affiliation: Work performed during an internship at Microsoft Project Leadhttps://llava-vl.github.io/llava-plus/
Abstract
This paper presents LLaVA-Plus (Large Language and Vision Assistants that Plug and Learn to Use Skills), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). LLaVA-Plus maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and is able to activate relevant tools, given users’ multimodal inputs, to compose their execution results on the fly to fulfill many real-world tasks. To acquire the ability of using tools, LLaVA-Plus is trained on multimodal instruction-following data that we have curated. The training data covers many tool use examples of visual understanding, generation, external knowledge retrieval and their compositions. Empirical results show that LLaVA-Plus outperforms LLaVA in existing capabilities, and exhibits many new capabilities. Compared with tool-augmented LLMs, LLaVA-Plus is distinct in that the image query is directly grounded in and actively engaged throughout the entire human-AI interaction sessions, significantly improving tool use performance and enabling new scenarios.
原文 arXiv:2311.05437;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2311.05437v1