Sociotechnical Safety Evaluation of Generative AI Systems
Laura Weidinger Google DeepMind, London N1C 4DN, United Kingdom Maribeth Rauh Google DeepMind, London N1C 4DN, United Kingdom Nahema Marchal Google DeepMind, London N1C 4DN, United Kingdom Arianna Manzini Google DeepMind, London N1C 4DN, United Kingdom Lisa Anne Hendricks Google DeepMind, London N1C 4DN, United Kingdom Juan Mateos-Garcia Google DeepMind, London N1C 4DN, United Kingdom Stevie Bergman Google DeepMind, London N1C 4DN, United Kingdom Jackie Kay Google DeepMind, London N1C 4DN, United Kingdom Conor Griffin Google DeepMind, London N1C 4DN, United Kingdom Ben Bariach Google DeepMind, London N1C 4DN, United Kingdom Iason Gabriel Google DeepMind, London N1C 4DN, United Kingdom Verena Rieser Google DeepMind, London N1C 4DN, United Kingdom William Isaac Google DeepMind, London N1C 4DN, United Kingdom
Abstract
Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a three-layered framework that takes a structured, sociotechnical approach to evaluating these risks. This framework encompasses capability evaluations, which are the main current approach to safety evaluation. It then reaches further by building on system safety principles, particularly the insight that context determines whether a given capability may cause harm. To account for relevant context, our framework adds human interaction and systemic impacts as additional layers of evaluation. Second, we survey the current state of safety evaluation of generative AI systems and create a repository of existing evaluations. Three salient evaluation gaps emerge from this analysis. We propose ways forward to closing these gaps, outlining practical steps as well as roles and responsibilities for different actors. Sociotechnical safety evaluation is a tractable approach to the robust and comprehensive safety evaluation of generative AI systems.
中文速览
生成式AI(Generative AI)系统正在各行各业快速普及,但随之而来的安全风险却缺乏系统、可比的评估方法。这篇文章提出了一套"社会技术安全评估框架",将评估分为三个递进的层次:能力评估(AI系统本身能做什么)、人机交互评估(真实用户在特定场景下如何与系统互动)、以及系统影响评估(AI在更大社会层面可能造成哪些连锁效应),核心思路是"语境决定风险",单靠测能力远远不够。作者同时对现有安全评估的实践现状做了全面调研,整理了一个公开的评估资源库,发现当前在人机交互和系统影响这两个层面存在明显的评估空白,并就如何填补这些空白给出了具体行动建议,以及AI开发者、监管者、公众政策制定者等各方应承担的责任。这项工作对于推动生成式AI安全评估走向标准化、综合化具有重要的参考价值,既能帮助开发者更系统地排查风险,也为政策制定者提供了清晰的行动框架。
原文 arXiv:2310.11986;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.11986v2