A Survey on LLM-as-a-Judge
Jiawei Gu1,2*, Xuhui Jiang1,3*, Zhichao Shi1,4,*, Hexiang Tan4, Xuehao Zhai5, Chengjin Xu1,3, Wei Li4, Yinghan Shen4, Shengjie Ma1,6, Honghao Liu1, Saizhuo Wang1,7,Kun Zhang4, Zhouchi Lin1, Bowen Zhang1, Lionel Ni7,8, Wen Gao9, Yuanzhuo Wang4,†, Jian Guo1,† Affiliation: 1IDEA Research, International Digital Economy AcademyChina 2Sun Yat-sen University , China 3DataArc Tech Ltd, China 4Institute of Computing Technology, Chinese Academy of Sciences , China 5Department of Civil and Environmental Engineering, Imperial College London, UK 6Gaoling School of Artificial Intelligence, Renmin University of China 7The Hong Kong University of Science and Technology , China 8The Hong Kong University of Science and Technology (Guangzhou) , China 9Department of Computer Science and Technology, Peking University , China
Abstract
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success across diverse domains, leading to the emergence of "LLM-as-a-Judge," where LLMs are employed as evaluators for complex tasks. With their ability to process diverse data types and provide scalable and flexible assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge that requires careful design and standardization. This paper provides a comprehensive survey on LLM-as-a-Judge, offering a formal definition and a detailed classification, while focusing on addressing the core question: How to built reliable LLM-as-a-Judge systems? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse assessment scenarios. Additionally, we propose methodologies for evaluating the reliability of LLM-as-a-Judge systems, supported by a novel benchmark designed for this purpose.
原文 arXiv:2411.15594;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2411.15594v6