A Survey on LLM-as-a-Judge
Jiawei Gu1,2*, Xuhui Jiang1,3*, Zhichao Shi1,4,*, Hexiang Tan4, Xuehao Zhai5, Chengjin Xu1,3, Wei Li4, Yinghan Shen4, Shengjie Ma1,6, Honghao Liu1, Saizhuo Wang1,7,Kun Zhang4, Zhouchi Lin1, Bowen Zhang1, Lionel Ni7,8, Wen Gao9, Yuanzhuo Wang4,†, Jian Guo1,† 1IDEA Research, International Digital Economy AcademyChina 2Sun Yat-sen University China 3DataArc Tech LtdChina 4Institute of Computing Technology, Chinese Academy of Sciences China 5Department of Civil and Environmental Engineering, Imperial College LondonUK 6Gaoling School of Artificial Intelligence, Renmin University of China 7The Hong Kong University of Science and Technology China 8The Hong Kong University of Science and Technology (Guangzhou) China 9Department of Computer Science and Technology, Peking University China
Abstract
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success across diverse domains, leading to the emergence of ”LLM-as-a-Judge,” where LLMs are employed as evaluators for complex tasks. With their ability to process diverse data types and provide scalable and flexible assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge that requires careful design and standardization. This paper provides a comprehensive survey on LLM-as-a-Judge, offering a formal definition and a detailed classification, while focusing on addressing the core question: How to built reliable LLM-as-a-Judge systems? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse assessment scenarios. Additionally, we propose methodologies for evaluating the reliability of LLM-as-a-Judge systems, supported by a novel benchmark designed for this purpose.
中文速览
用大语言模型(LLM)来充当评审员(LLM-as-a-Judge)已成为替代人工专家评估的新兴范式,但如何保证其评判结果的可靠性至今缺乏系统性研究。这篇综述给出了 LLM-as-a-Judge 的正式定义与分类框架,并围绕"如何构建可靠的 LLM 评审系统"这一核心问题,系统梳理了提升一致性、消除偏见、适应多样评估场景的策略,同时提出了一套专门用于衡量系统可靠性的评测基准。研究结果表明,LLM 评审在扩展性和上下文理解能力上兼具传统自动指标与人工专家评估的优势,但仍面临元评估、长期一致性等尚未解决的挑战。这项工作为学术界和工业界提供了从概念界定到实践部署的完整参考,有助于推动构建值得社会信赖的 AI 评审系统。
原文 arXiv:2411.15594;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2411.15594v6