Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper MIT CSAIL Xander Davies Harvard UniversityClaudia Shi, Columbia UniversityThomas Krendl Gilbert, Cornell TechJérémy Scheurer, Apollo ResearchJavier Rando, ETH ZurichRachel Freedman, UC BerkeleyTomasz Korbak, University of SussexDavid Lindner, ETH ZurichPedro Freire, IndependentTony Wang, MIT CSAILSamuel Marks, Harvard UniversityCharbel-Raphaël Segerie, EffiSciencesMicah Carroll, UC BerkeleyAndi Peng, MIT CSAILPhillip Christoffersen, MIT CSAILMehul Damani, MIT CSAILStewart Slocum, MIT CSAILUsman Anwar, University of CambridgeAnand Siththaranjan, UC BerkeleyMax Nadeau, Harvard UniversityEric J. Michaud, MITJacob Pfau, New York UniversityDmitrii Krasheninnikov, University of CambridgeXin Chen, ETH ZurichLauro Langosco, University of CambridgePeter Hase, UNC Chapel HillErdem Bıyık, University of Southern CaliforniaAnca Dragan, UC BerkeleyDavid Krueger, University of CambridgeDorsa Sadigh, Stanford UniversityDylan Hadfield-Menell, MIT CSAIL
Abstract
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of-the-art large language models (LLMs). Despite this popularity, there has been relatively little public work systematizing its flaws. In this paper, we (1) survey open problems and fundamental limitations of RLHF and related methods; (2) overview techniques to understand, improve, and complement RLHF in practice; and (3) propose auditing and disclosure standards to improve societal oversight of RLHF systems. Our work emphasizes the limitations of RLHF and highlights the importance of a multi-layered approach to the development of safer AI systems.
原文 arXiv:2307.15217;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2307.15217v2