Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models
Maribeth Rauh、John Mellor、Jonathan Uesato、Po-Sen Huang、Johannes Welbl、Laura Weidinger、Sumanth Dathathri、Amelia Glaese、Geoffrey Irving、Iason Gabriel、William Isaac、Lisa Anne Hendricks \ANDDeepMind Corresponding author:
Abstract
Large language models produce human-like text that drives a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic, biased, untruthful or otherwise harmful. Though work to evaluate language model harms is under way, translating foresight about which harms may arise into rigorous benchmarks is not straightforward. To facilitate this translation, we outline six ways of characterizing harmful text which merit explicit consideration when designing new benchmarks. We then use these characteristics as a lens to identify trends and gaps in existing benchmarks. Finally, we apply them in a case study of the Perspective API, a toxicity classifier that is widely used in harm benchmarks. Our characteristics provide one piece of the bridge that translates between foresight and effective evaluation.
中文速览
大语言模型(Large Language Model, LLM)生成的文字越来越像人类写作,但同时也可能输出带有毒性、偏见或虚假信息的有害内容,如何系统地评估这些危害却远非易事。作者提出了六个刻画有害文本的关键维度——危害定义、表征/分配/能力公平性、个例与分布层面的危害、文本/应用/社会背景、危害承受方,以及涉及的人口群体——为基准测试(benchmark)的设计提供了一套共同语言。通过将现有基准映射到这六个维度,研究者发现了明显的评估盲区,例如几乎没有基准考察长文本语境下的危害;并以广泛使用的毒性分类器 Perspective API 为案例,揭示了其设计初衷与实际用途之间的错位。这项工作为连接"预见潜在危害"与"严格量化评估"之间的鸿沟搭建了一座概念桥梁,有助于推动未来更完善、更负责任的语言模型危害评估体系的建立。
原文 arXiv:2206.08325;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.08325v2