Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Yancheng He∗, Shilong Li∗, Jiaheng Liu∗†, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Xuepeng Liu, Dekai Sun, Shirong Lin, Zhicheng Zheng, Xiaoyong Zhu, Wenbo Su, Bo Zheng Taobao、Tmall Group of Alibaba
Abstract
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models.
中文速览
大型语言模型(LLM)在中文场景下的事实性知识掌握程度缺乏专门的评测工具,现有基准要么以英文为主,要么难以高效打分。为此,研究团队构建了"Chinese SimpleQA"——首个专注于中文事实性问答的综合基准,包含3000道横跨中国文化、人文、理工、生活艺术、社会与自然科学六大主题的简短问答题,并通过自动化生成加严格人工审核的双重流程确保质量,所有参考答案固定不变、随时间不过期,且可直接调用LLM API快速评分。在对40余个主流闭源与开源模型的系统评测中,发现多数模型得分偏低,仅o1-preview和豆包pro-32k勉强及格,中文社区模型在"中国文化"子话题上显著优于GPT/o1系列,而模型越大、校准性越好、事实准确率也越高,引入检索增强生成(RAG)后模型间差距大幅收窄。该基准为开发者提供了一把衡量模型中文事实能力的标尺,有助于推动基础模型在中文知识层面的持续改进。
原文 arXiv:2411.07140;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2411.07140v2