arXiv:2210.10760 · 中英对照阅读
Scaling Laws for Reward Model Overoptimization
中文速览
研究关注强化学习从人类反馈中常见的“代理奖励被过度优化”问题:模型越迎合奖励模型,真实偏好反而可能变差。作者用一个大型“黄金”奖励模型模拟人类偏好,再训练不同规模的代理奖励模型,并分别用强化学习和最佳候选采样(best-of-n)不断优化它,观察真实奖励如何变化。结果显示,两种优化方式都先提升、后因“古德哈特定律”出现过度优化,而且真实奖励与优化程度之间存在可预测的函数关系,其参数会随奖励模型规模平滑变化;更多训练数据能减少问题,策略模型大小影响收益却不明显改变过度优化程度,KL惩罚也未带来清晰的真实收益改善。研究的重要性在于,它用低成本的合成实验定量刻画了奖励模型何时会失效,为理解和控制RLHF中的代理目标偏移及未来AI对齐风险提供了可检验的经验规律。
摘要
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder ground truth performance, in accordance with Goodhart’s law. This effect has been frequently observed, but not carefully measured due to the expense of collecting human preference data. In this work, we use a synthetic setup in which a fixed “gold-standard” reward model plays the role of humans, providing labels used to train a proxy reward model. We study how the gold reward model score changes as we optimize against the proxy reward model using either reinforcement learning or best-of- $n$ sampling. We find that this relationship follows a different functional form depending on the method of optimization, and that in both cases its coefficients scale smoothly with the number of reward model parameters. We also study the effect on this relationship of the size of the reward model dataset, the number of reward model and policy parameters, and the coefficient of the KL penalty added to the reward in the reinforcement learning setup. We explore the implications of these empirical results for theoretical considerations in AI alignment.
术语表
- Reinforcement Learning from Human Feedback (RLHF)
- 基于人类反馈的强化学习(RLHF)
- reward model (RM)
- 奖励模型(RM)
- proxy reward model
- 代理奖励模型
- gold-standard reward model
- 金标准奖励模型
- Goodhart’s law
- 古德哈特定律
- overoptimization
- 过度优化
- human preference
- 人类偏好
- best-of-n (BoN) sampling
- 最佳候选采样(BoN)
- rejection sampling
- 拒绝采样
- reranking
- 重排序
- policy gradient
- 策略梯度
- reinforcement learning (RL)
- 强化学习(RL)
- Kullback–Leibler divergence (KL divergence)
- Kullback–Leibler 散度(KL 散度)
- KL penalty
- KL 惩罚项
- scaling law
- 缩放定律
- AI alignment
- AI 对齐
- Proximal Policy Optimization (PPO)
- 近端策略优化(PPO)
- supervised fine-tuning (SFT)
- 监督微调(SFT)
- policy
- 策略
- initial policy
- 初始策略
- proxy objective
- 代理目标
- ground truth
- 真实值
- synthetic data
- 合成数据
- synthetic setup
- 合成设置
- preference labels
- 偏好标签
- language model
- 语言模型
- GPT-3
- GPT-3
- scalar head
- 标量头
- rollout
- 轨迹采样