MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu1,2,♠, Peixian Chen3, Yunhang Shen3, Yulei Qin3, Mengdan Zhang3 Xu Lin3, Jinrui Yang3, Xiawu Zheng4, Ke Li3,†, Xing Sun3 Yunsheng Wu3, Rongrong Ji4, Caifeng Shan1,2, Ran He5 1State Key Laboratory for Novel Software Technology, Nanjing University 2School of Intelligence Science and Technology, Nanjing University 3Tencent Youtu Lab 4Xiamen University 5CASIA ♠ Project Leader † Corresponding Author
Abstract
Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page: https://github.com/BradyFU/Awesome-Multimodal-Large-L
中文速览
多模态大语言模型(Multimodal Large Language Model,MLLM)虽然能写诗、做推理,展现出令人惊艳的涌现能力,但学界一直缺乏一套系统、公平的量化评测基准来摸清它们到底"有多强"。为此,研究团队构建了首个全面的MLLM评测基准MME,涵盖感知(如物体存在、数量、颜色、位置以及名人、地标等细粒度识别)和认知(如常识推理、数值计算、文字翻译、代码理解)共14个子任务,所有题目的图文问答对均由人工标注,以避免训练集数据泄露,并统一采用"请回答是或否"的简洁指令设计,让量化统计既客观又公平。他们在MME上系统评测了30个主流MLLM的零样本表现,发现各模型差距显著,且普遍存在无法遵循基本指令、感知与推理能力薄弱、物体幻觉(object hallucination)等突出问题。这项工作填补了MLLM综合评测的空白,为后续模型优化指明了方向,具有重要的参考价值。
原文 arXiv:2306.13394;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.13394v5