Detecting Hate Speech with GPT-3 Thanks: Code and data are available at: https://github.com/kelichiu/GPT3-hate-speech-detection. We gratefully acknowledge the support of Gillian Hadfield, the Schwartz Reisman Institute for Technology and Society, and OpenAI for providing access to GPT-3 under the academic access program. We thank two anonymous reviews and the editor, as well as Amy Farrow, Christina Nguyen, Haoluan Chen, John Giorgi, Mauricio Vargas Sepúlveda, Monica Alexander, Noam Kolt, and Tom Davidson for helpful discussions and suggestions. Please note that we have added asterisks to racial slurs and other offensive content in this paper, however the inputs and outputs did not have these. Comments on the 24 March 2022 version of this paper are welcome at: rohan.alexander@utoronto.ca.
Ke-Li Chiu University of Toronto Annie Collins University of Toronto Rohan Alexander University of Toronto and Schwartz Reisman Institute
Abstract
Sophisticated language models such as OpenAI’s GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech and classify text as sexist or racist. We use GPT-3 to identify sexist and racist text passages with zero-, one-, and few-shot learning. We find that with zero- and one-shot learning, GPT-3 can identify sexist or racist text with an average accuracy between 55 per cent and 67 per cent, depending on the category of text and type of learning. With few-shot learning, the model’s accuracy can be as high as 85 per cent. Large language models have a role to play in hate speech detection, and with further development they could eventually be used to counter hate speech.
原文 arXiv:2103.12407;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.12407v4