A StrongREJECT¯¯StrongREJECT\overline{\hbox{{StrongREJECT}}}over¯ start_ARG StrongREJECT end_ARG for Empty Jailbreaks
Alexandra Souly∗、Qingyuan Lu∗、Dillon Bowen∗ Tu Trinh†、Elvis Hsieh†、Sana Pandey、Pieter Abbeel、Justin Svegliato Scott Emmons‡、Olivia Watkins‡、Sam Toyer‡ Center for Human-Compatible AI, UC Berkeley
Abstract
Most jailbreak papers claim the jailbreaks they propose are highly effective, often boasting near-100% attack success rates. However, it is perhaps more common than not for jailbreak developers to substantially exaggerate the effectiveness of their jailbreaks. We suggest this problem arises because jailbreak researchers lack a standard, high-quality benchmark for evaluating jailbreak performance, leaving researchers to create their own. To create a benchmark, researchers must choose a dataset of forbidden prompts to which a victim model will respond, along with an evaluation method that scores the harmfulness of the victim model’s responses. We show that existing benchmarks suffer from significant shortcomings and introduce the StrongREJECT benchmark to address these issues. StrongREJECT’s dataset contains prompts that victim models must answer with specific, harmful information, while its automated evaluator measures the extent to which a response gives useful information to forbidden prompts. In doing so, the StrongREJECT evaluator achieves state-of-the-art agreement with human judgments of jailbreak effectiveness. Notably, we find that existing evaluation methods significantly o
中文速览
越狱攻击(jailbreak)研究领域长期存在一个严重问题:大量论文声称自己的攻击方法成功率接近100%,但这些数字往往严重虚高,因为各家研究者自行设计的评测基准质量参差不齐,既有问题模糊、答案无法验证的提示词,也有只要模型"没有拒绝"就算攻击成功的粗糙评分标准。为此,研究团队构建了StrongREJECT基准,包含313条经过严格筛选的禁止性提示词(覆盖非法物品、仇恨歧视、暴力等六大类),以及一个自动评分器——该评分器不仅判断模型是否拒绝回答,还同时衡量回答内容是否真正具体、有用,从而与人工评判高度一致,达到了同类评估方法的最优水平。实验结果显示,现有评测方法普遍高估了越狱效果,根本原因在于一个此前未被注意的现象:越狱手段在绕过安全训练的同时,往往会同步削弱模型的实际能力,导致输出内容空洞无用。这项工作为整个越狱研究领域提供了一把更可靠的量尺,有助于研究者准确识别哪些攻击真正构成威胁,推动安全评测走向标准化。
原文 arXiv:2402.10260;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2402.10260v2