Sociotechnical Safety Evaluation of Generative AI Systems
Laura Weidinger Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Maribeth Rauh Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Nahema Marchal Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Arianna Manzini Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Lisa Anne Hendricks Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Juan Mateos-Garcia Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Stevie Bergman Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Jackie Kay Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Conor Griffin Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Ben Bariach Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Iason Gabriel Affiliation: Google DeepMind, London N1C 4DN, United Kingdom Verena Rieser Affiliation: Google DeepMind, London N1C 4DN, United Kingdom William Isaac Affiliation: Google DeepMind, London N1C 4DN, United Kingdom
Abstract
Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a three-layered framework that takes a structured, sociotechnical approach to evaluating these risks. This framework encompasses capability evaluations, which are the main current approach to safety evaluation. It then reaches further by building on system safety principles, particularly the insight that context determines whether a given capability may cause harm. To account for relevant context, our framework adds human interaction and systemic impacts as additional layers of evaluation. Second, we survey the current state of safety evaluation of generative AI systems and create a repository of existing evaluations. Three salient evaluation gaps emerge from this analysis. We propose ways forward to closing these gaps, outlining practical steps as well as roles and responsibilities for different actors. Sociotechnical safety evaluation is a tractable approach to the robust and comprehensive safety evaluation of generative AI systems.
原文 arXiv:2310.11986;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.11986v2