ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
Petter Törnberg Amsterdam Institute for Social Science Research (AISSR), University of Amsterdam
Abstract
This paper assesses the accuracy, reliability and bias of the Large Language Model (LLM) ChatGPT-4 on the text analysis task of classifying the political affiliation of a Twitter poster based on the content of a tweet. The LLM is compared to manual annotation by both expert classifiers and crowd workers, generally considered the gold standard for such tasks. We use Twitter messages from United States politicians during the 2020 election, providing a ground truth against which to measure accuracy. The paper finds that ChatGPT-4 has achieves higher accuracy, higher reliability, and equal or lower bias than the human classifiers. The LLM is able to correctly annotate messages that require reasoning on the basis of contextual knowledge, and inferences around the author’s intentions – traditionally seen as uniquely human abilities. These findings suggest that LLM will have substantial impact on the use of textual data in the social sciences, by enabling interpretive research at a scale.
中文速览
推特上的一条帖子,究竟能不能从内容判断出发帖者是民主党人还是共和党人?这正是本文要解答的问题。研究者用2020年美国大选前两个月美国参议员发布的500条推文作为测试集(因为议员党籍已知,可以客观验证准确率),分别让ChatGPT-4、众包平台MTurk工人和政治学专家来做判断,并从准确率、一致性和偏差三个维度横向比较。结果发现,ChatGPT-4在三项指标上均优于或持平于人类标注者——它不仅能识别直白的党派信号,还能处理需要背景知识推断、乃至揣摩作者意图的复杂推文,而这类能力向来被认为是人类独有的。这一发现意味着,大型语言模型(LLM)有望让社会科学研究者在保持解释性深度的同时大幅扩展分析规模,从根本上改变文本数据的使用方式。
原文 arXiv:2304.06588;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2304.06588v1