How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
Biyang Guo1† , Xin Zhang2∗, Ziyuan Wang1∗, Minqi Jiang1∗, Jinran Nie3∗ Yuxuan Ding4, Jianwei Yue5, Yupeng Wu6 1AI Lab, School of Information Management and Engineering Shanghai University of Finance and Economics 2Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen) 3School of Information Science, Beijing Language and Culture University 4School of Electronic Engineering, Xidian University 5School of Computing, Queen’s University, 6Wind Information Co., Ltd Equal Contribution.
Abstract
The introduction of ChatGPT111Launched by OpenAI in November 2022. https://chat.openai.com/chat has garnered widespread attention in both academic and industrial communities. ChatGPT is able to respond effectively to a wide range of human questions, providing fluent and comprehensive answers that significantly surpass previous public chatbots in terms of security and usefulness. On one hand, people are curious about how ChatGPT is able to achieve such strength and how far it is from human experts. On the other hand, people are starting to worry about the potential negative impacts that large language models (LLMs) like ChatGPT could have on society, such as fake news, plagiarism, and social security issues. In this work, we collected tens of thousands of comparison responses from both human experts and ChatGPT, with questions ranging from open-domain, financial, medical, legal, and psychological areas. We call the collected dataset the Human ChatGPT Comparison Corpus (HC3). Based on the HC3 dataset, we study the characteristics of ChatGPT’s responses, the differences and gaps from human experts, and future directions for LLMs. We conducted comprehensive human evaluations and lingui
中文速览
大规模语言模型ChatGPT的横空出世让人们既兴奋又担忧——它的回答到底和真人专家差多少,又该如何识别AI生成的内容?为此,研究者构建了一个名为"人类-ChatGPT对比语料库"(HC3,Human ChatGPT Comparison Corpus)的数据集,收录了近4万组问题及人类专家与ChatGPT分别给出的回答,涵盖开放域、金融、医疗、法律、心理等多个领域。基于这份数据,研究团队进行了系统性的人工评测和语言学分析,发现ChatGPT的回答更有条理、篇幅更长、偏见更少,但在专业领域存在"一本正经地编造事实"的风险,且超过半数情况下被用户认为比人类回答更有帮助。在此基础上,他们还训练了多个自动检测模型,能够有效判断一段文字究竟出自人类还是ChatGPT,为平台治理AI生成内容提供了实用工具。这项工作首次系统地从语言特征和可检测性两个维度对ChatGPT与人类的差异进行了量化研究,数据集、代码与模型均已开源,对学术界和内容平台的监管都具有重要参考价值。
原文 arXiv:2301.07597;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2301.07597v1