Aha.
正在载入中英对照阅读…

arXiv:2210.10760 · 中英对照阅读

Scaling Laws for Reward Model Overoptimization

Leo Gao、John Schulman、Jacob Hilton

中文速览

研究关注强化学习从人类反馈中常见的“代理奖励被过度优化”问题:模型越迎合奖励模型,真实偏好反而可能变差。作者用一个大型“黄金”奖励模型模拟人类偏好,再训练不同规模的代理奖励模型,并分别用强化学习和最佳候选采样(best-of-n)不断优化它,观察真实奖励如何变化。结果显示,两种优化方式都先提升、后因“古德哈特定律”出现过度优化,而且真实奖励与优化程度之间存在可预测的函数关系,其参数会随奖励模型规模平滑变化;更多训练数据能减少问题,策略模型大小影响收益却不明显改变过度优化程度,KL惩罚也未带来清晰的真实收益改善。研究的重要性在于,它用低成本的合成实验定量刻画了奖励模型何时会失效,为理解和控制RLHF中的代理目标偏移及未来AI对齐风险提供了可检验的经验规律。

摘要

In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder ground truth performance, in accordance with Goodhart’s law. This effect has been frequently observed, but not carefully measured due to the expense of collecting human preference data. In this work, we use a synthetic setup in which a fixed “gold-standard” reward model plays the role of humans, providing labels used to train a proxy reward model. We study how the gold reward model score changes as we optimize against the proxy reward model using either reinforcement learning or best-of- $n$ sampling. We find that this relationship follows a different functional form depending on the method of optimization, and that in both cases its coefficients scale smoothly with the number of reward model parameters. We also study the effect on this relationship of the size of the reward model dataset, the number of reward model and policy parameters, and the coefficient of the KL penalty added to the reward in the reinforcement learning setup. We explore the implications of these empirical results for theoretical considerations in AI alignment.

术语表

Reinforcement Learning from Human Feedback (RLHF)
基于人类反馈的强化学习(RLHF)
reward model (RM)
奖励模型(RM)
proxy reward model
代理奖励模型
gold-standard reward model
金标准奖励模型
Goodhart’s law
古德哈特定律
overoptimization
过度优化
human preference
人类偏好
best-of-n (BoN) sampling
最佳候选采样(BoN)
rejection sampling
拒绝采样
reranking
重排序
policy gradient
策略梯度
reinforcement learning (RL)
强化学习(RL)
Kullback–Leibler divergence (KL divergence)
Kullback–Leibler 散度(KL 散度)
KL penalty
KL 惩罚项
scaling law
缩放定律
AI alignment
AI 对齐
Proximal Policy Optimization (PPO)
近端策略优化(PPO)
supervised fine-tuning (SFT)
监督微调(SFT)
policy
策略
initial policy
初始策略
proxy objective
代理目标
ground truth
真实值
synthetic data
合成数据
synthetic setup
合成设置
preference labels
偏好标签
language model
语言模型
GPT-3
GPT-3
scalar head
标量头
rollout
轨迹采样