Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Timothy R. McIntosh\orcidlink0000-0003-0836-4266*, Teo Susnjak\orcidlink0000-0001-9416-1435, Nalin Arachchilage\orcidlink0000-0002-0059-0376, Tong Liu\orcidlink0000-0003-3047-1148, Dan Xu\orcidlink0009-0004-3930-7381, Paul Watters\orcidlink0000-0002-1399-7175, , and Malka N. Halgamuge\orcidlink0000-0001-9994-3778 Manuscript received October 14, 2024. Corresponding Author: Timothy R. McIntosh (e-mail:
Abstract
The rapid rise in popularity of Large Language Models (LLMs) with emerging capabilities has spurred public curiosity to evaluate and compare different LLMs, leading many researchers to propose their own LLM benchmarks. Noticing preliminary inadequacies in those benchmarks, we embarked on a study to critically assess 23 state-of-the-art LLM benchmarks, using our novel unified evaluation framework through the lenses of people, process, and technology, under the pillars of benchmark functionality and integrity. Our research uncovered significant limitations, including biases, difficulties in measuring genuine reasoning, adaptability, implementation inconsistencies, prompt engineering complexity, evaluator diversity, and the overlooking of cultural and ideological norms in one comprehensive assessment. Our discussions emphasized the urgent need for standardized methodologies, regulatory certainties, and ethical guidelines in light of Artificial Intelligence (AI) advancements, including advocating for an evolution from static benchmarks to dynamic behavioral profiling to accurately capture LLMs’ complex behaviors and potential risks. Our study highlighted the necessity for a paradigm sh
中文速览
当前大语言模型(LLM)领域基准测试泛滥成灾,各家研究者纷纷推出自己的评测方案,却普遍存在偏见、无法真正衡量推理能力、忽视文化差异、容易被模型"刷分"等系统性缺陷。为此,研究团队提出了一套统一评估框架,从"人、流程、技术"三个维度出发,兼顾基准的功能性与完整性,对23个主流LLM基准进行了全面批判性审查。结果发现,几乎所有基准都在不同程度上存在评估指标片面、提示工程复杂度失控、评测者多样性不足、忽略意识形态与文化规范等问题。研究进一步倡导将静态的考试式基准升级为动态的行为画像(behavioral profiling)与定期审计机制,以更真实地捕捉LLM在现实场景中的能力边界与潜在风险。这项工作的意义在于,它为行业建立标准化方法论、监管框架和伦理准则提供了系统性依据,推动LLM评测范式从"比分数"转向"测行为",从而更负责任地将AI系统融入社会。
原文 arXiv:2402.09880;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2402.09880v2