RL with KL penalties is better viewed as Bayesian inference
Tomasz Korbak Affiliation: University of Sussex Affiliation: New York University Email: Ethan Perez Affiliation: New York University Email: Christopher L Buckley Affiliation: University of Sussex Email:
Abstract
Reinforcement learning (RL) is frequently employed in fine-tuning large language models (LMs) to penalize them for undesirable features of generated sequences, such as offensiveness or harmfulness. In this paper, we analyze challenges associated with treating language models as RL policies and show how avoiding those challenges requires moving beyond the RL paradigm.
原文 arXiv:2205.11275;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.11275v2