On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
Jindong Wang1 Contact: 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Xixu Hu1,2‡ Equal contribution. 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Wenxin Hou3† 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Hao Chen4 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Runkai Zheng1,5 Work done during internship at Microsoft Research Asia. 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Yidong Wang6 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Linyi Yang7 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Wei Ye6 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Haojun Huang3 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Xiubo Geng3 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Binxing Jiao3 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Yue Zhang7 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn Xing Xie1 1Microsoft Research, 2City University of Hong Kong, 3Microsoft STCA, 4Carnegie Mellon University, 5Chinese University of Hong Kong (Shenzhen), 6Peking University, 7Westlake University https://github.com/microsoft/robustlearn
Abstract
ChatGPT is a recent chatbot service released by OpenAI and is receiving increasing attention over the past few months. While evaluations of various aspects of ChatGPT have been done, its robustness, i.e., the performance to unexpected inputs, is still unclear to the public. Robustness is of particular concern in responsible AI, especially for safety-critical applications. In this paper, we conduct a thorough evaluation of the robustness of ChatGPT from the adversarial and out-of-distribution (OOD) perspective. To do so, we employ the AdvGLUE and ANLI benchmarks to assess adversarial robustness and the Flipkart review and DDXPlus medical diagnosis datasets for OOD evaluation. We select several popular foundation models as baselines. Results show that ChatGPT shows consistent advantages on most adversarial and OOD classification and translation tasks. However, the absolute performance is far from perfection, which suggests that adversarial and OOD robustness remains a significant threat to foundation models. Moreover, ChatGPT shows astounding performance in understanding dialogue-related texts and we find that it tends to provide informal suggestions for medical tasks instead of defi
中文速览
ChatGPT凭借庞大的训练数据和强大的对话能力迅速风靡全球,但它在面对"刁钻输入"时是否依然可靠,一直缺乏系统性验证。研究者从对抗鲁棒性(adversarial robustness)和分布外鲁棒性(out-of-distribution robustness)两个维度,用AdvGLUE、ANLI、Flipkart评论及DDXPlus医疗诊断等多个基准,对ChatGPT与多个主流大语言模型进行了零样本(zero-shot)横向比较。结果显示,ChatGPT在大多数对抗性和分布外分类任务上确实优于其他对手,在理解对话文本和机器翻译方面表现尤为突出,但绝对准确率仍远未达到可靠部署的水准,说明鲁棒性缺陷依然是所有大模型落地安全关键场景的重大隐患。这项研究为业界评估和改进大语言模型的可靠性提供了系统性证据,也为未来如何在零样本条件下强化模型鲁棒性指出了研究方向。
原文 arXiv:2302.12095;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2302.12095v5