How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
Biyang Guo Thanks: Equal Contribution. Affiliation: AI Lab, School of Information Management and EngineeringShanghai University of Finance and Economics Xin Zhang Affiliation: Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen) Ziyuan Wang Affiliation: School of Information Science, Beijing Language and Culture University Minqi Jiang Jinran Nie Yuxuan Ding Affiliation: School of Electronic Engineering, Xidian University Jianwei Yue Yupeng Wu Affiliation: School of Computing, Queen’s University, Wind Information Co., Ltd
Abstract
The introduction of ChatGPT11 1 Launched by OpenAI in November 2022. https://chat.openai.com/chat has garnered widespread attention in both academic and industrial communities. ChatGPT is able to respond effectively to a wide range of human questions, providing fluent and comprehensive answers that significantly surpass previous public chatbots in terms of security and usefulness. On one hand, people are curious about how ChatGPT is able to achieve such strength and how far it is from human experts. On the other hand, people are starting to worry about the potential negative impacts that large language models (LLMs) like ChatGPT could have on society, such as fake news, plagiarism, and social security issues. In this work, we collected tens of thousands of comparison responses from both human experts and ChatGPT, with questions ranging from open-domain, financial, medical, legal, and psychological areas. We call the collected dataset the Human ChatGPT Comparison Corpus (HC3). Based on the HC3 dataset, we study the characteristics of ChatGPT’s responses, the differences and gaps from human experts, and future directions for LLMs. We conducted comprehensive human evaluations and ling
原文 arXiv:2301.07597;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2301.07597v1