Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Rajkumar Ramamurthy Prithviraj Ammanabrolu Kianté Brantley Affiliation: Cornell University Jack Hessel Affiliation: Allen Institute for Artificial Intelligence Rafet Sifa Affiliation: Fraunhofer IAIS Christian Bauckhage Affiliation: Fraunhofer IAIS Hannaneh Hajishirzi Yejin Choi
Abstract
We tackle the problem of aligning pre-trained large language models (LMs) with human preferences. If we view text generation as a sequential decision-making problem, reinforcement learning (RL) appears to be a natural conceptual framework. However, using RL for LM-based generation faces empirical challenges, including training instability due to the combinatorial action space, as well as a lack of open-source libraries and benchmarks customized for LM alignment. Thus, a question rises in the research community: is RL a practical paradigm for NLP?
原文 arXiv:2210.01241;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2210.01241v3