Automatic Detection of Machine Generated Text: A Critical Survey
Ganesh Jawahar, Muhammad Abdul-Mageed, Laks V.S. Lakshmanan University of British Columbia, Vancouver, Canada
Abstract
Text generative models (TGMs) excel in producing text that matches the style of human language reasonably well. Such TGMs can be misused by adversaries, e.g., by automatically generating fake news and fake product reviews that can look authentic and fool humans. Detectors that can distinguish text generated by TGM from human written text play a vital role in mitigating such misuse of TGMs. Recently, there has been a flurry of works from both natural language processing (NLP) and machine learning (ML) communities to build accurate detectors for English. Despite the importance of this problem, there is currently no work that surveys this fast-growing literature and introduces newcomers to important research challenges. In this work, we fill this void by providing a critical survey and review of this literature to facilitate a comprehensive understanding of this problem. We conduct an in-depth error analysis of the state-of-the-art detector and discuss research directions to guide future work in this exciting area.
中文速览
机器生成文本检测(machine-generated text detection)正随着GPT-2、GPT-3等大模型的滥用风险而变得日益紧迫,自动生成的假新闻、虚假评论已能以假乱真、骗过人眼。这篇文章对这一快速增长领域的研究进行了首次系统性综述,梳理了现有检测方法的原理与分类,并对当前最先进检测器进行了深入的错误分析,指出其在泛化性、鲁棒性和可解释性等方面的不足。研究发现,现有检测器在面对不同模型架构、解码策略和领域迁移时表现明显下降,难以应对真实场景中的对抗威胁。这份综述为入门者提供了清晰的知识框架,也为研究者指出了亟待突破的方向,对推动该领域走向更实用、更可靠的检测系统具有重要的参考价值。
原文 arXiv:2011.01314;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2011.01314v1