Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper,∗ MIT CSAIL, Xander Davies,∗ Harvard University Claudia Shi, Columbia University Thomas Krendl Gilbert, Cornell Tech Jérémy Scheurer, Apollo Research Javier Rando, ETH Zurich Rachel Freedman, UC Berkeley Tomasz Korbak, University of Sussex David Lindner, ETH Zurich Pedro Freire, Independent Tony Wang, MIT CSAIL Samuel Marks, Harvard University Charbel-Raphaël Segerie, EffiSciences Micah Carroll, UC Berkeley Andi Peng, MIT CSAIL Phillip Christoffersen, MIT CSAIL Mehul Damani, MIT CSAIL Stewart Slocum, MIT CSAIL Usman Anwar, University of Cambridge Anand Siththaranjan, UC Berkeley Max Nadeau, Harvard University Eric J. Michaud, MIT Jacob Pfau, New York University Dmitrii Krasheninnikov, University of Cambridge Xin Chen, ETH Zurich Lauro Langosco, University of Cambridge Peter Hase, UNC Chapel Hill Erdem Bıyık, University of Southern California Anca Dragan, UC Berkeley David Krueger, University of Cambridge Dorsa Sadigh, Stanford University Dylan Hadfield-Menell, MIT CSAIL
Abstract
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of-the-art large language models (LLMs). Despite this popularity, there has been relatively little public work systematizing its flaws. In this paper, we (1) survey open problems and fundamental limitations of RLHF and related methods; (2) overview techniques to understand, improve, and complement RLHF in practice; and (3) propose auditing and disclosure standards to improve societal oversight of RLHF systems. Our work emphasizes the limitations of RLHF and highlights the importance of a multi-layered approach to the development of safer AI systems.
中文速览
大语言模型(LLM)对齐领域最流行的方法——基于人类反馈的强化学习(RLHF,Reinforcement Learning from Human Feedback)——在GPT-4、Claude等顶尖模型中被广泛使用,但学界和业界对其缺陷却缺乏系统性的公开梳理。这篇论文对此做了全面的调查与分类:从人类标注者的偏见与恶意、奖励模型(reward model)的过度优化与分布偏移,到强化学习策略训练中的不稳定性与"讨好"行为,作者逐层剖析了RLHF在反馈收集、奖励建模和策略优化三个环节中的开放问题与根本局限。研究发现,RLHF并不能独立保障AI安全,泄露隐私、产生幻觉、放大政治偏见、难以抵御"越狱"攻击等问题依然普遍存在,因此需要可解释性工具、红队测试、形式化验证等多层技术手段共同补充,同时还需要建立针对RLHF系统的审计与信息披露规范。这项工作的重要性在于:它首次系统地为研究者、工程师和政策制定者勾勒出RLHF的风险图谱,呼吁各方共同推动更安全、更可问责的AI开发实践。
原文 arXiv:2307.15217;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2307.15217v2