“Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Xinyue Shen1 Zeyuan Chen1 Michael Backes1 Yun Shen2 Yang Zhang1 1CISPA Helmholtz Center for Information Security 2NetApp Yang Zhang is the corresponding author.
Abstract
The misuse of large language models (LLMs) has drawn significant attention from the general public and LLM vendors. One particular type of adversarial prompt, known as jailbreak prompt, has emerged as the main attack vector to bypass the safeguards and elicit harmful content from LLMs. In this paper, employing our new framework JailbreakHub, we conduct a comprehensive analysis of 1,405 jailbreak prompts spanning from December 2022 to December 2023. We identify 131 jailbreak communities and discover unique characteristics of jailbreak prompts and their major attack strategies, such as prompt injection and privilege escalation. We also observe that jailbreak prompts increasingly shift from online Web communities to prompt-aggregation websites and 28 user accounts have consistently optimized jailbreak prompts over 100 days. To assess the potential harm caused by jailbreak prompts, we create a question set comprising 107,250 samples across 13 forbidden scenarios. Leveraging this dataset, our experiments on six popular LLMs show that their safeguards cannot adequately defend jailbreak prompts in all scenarios. Particularly, we identify five highly effective jailbreak prompts that achiev
中文速览
大型语言模型(LLM)被"越狱提示词(jailbreak prompt)"绕过安全机制、生成有害内容的问题日益严峻,但学界此前缺乏对这一现象的系统性认识。研究者构建了名为 JailbreakHub 的分析框架,从 Reddit、Discord、网站和开源数据集等四类平台收集了 2022 年 12 月至 2023 年 12 月间的 1,405 条真实越狱提示词,识别出 131 个越狱社区,并揭示了提示注入、权限提升、虚拟化场景等主要攻击策略。为评估实际危害,研究者还构建了涵盖 13 类违禁场景、共 107,250 条问题的测试集,在 ChatGPT(GPT-3.5)、GPT-4 等六款主流模型上进行实验,发现最有效的越狱提示词攻击成功率高达 0.95,且某些提示词在网上持续流传超过 240 天,而现有内外部安全机制对其防御效果均十分有限。这项研究首次系统描绘了野生越狱提示词的全貌,为 LLM 厂商和监管方完善安全防线提供了重要的实证依据。
原文 arXiv:2308.03825;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2308.03825v2