Reward-Free Exploration for Reinforcement Learning
Chi Jin Affiliation: Princeton University Email: Akshay Krishnamurthy Affiliation: Microsoft Research, New York Email: Max Simchowitz Affiliation: University of California, Berkeley Email: Tiancheng Yu Affiliation: Massachusetts Institute of Technology Email:
Abstract
Exploration is widely regarded as one of the most challenging aspects of reinforcement learning (RL), with many naive approaches succumbing to exponential sample complexity. To isolate the challenges of exploration, we propose a new “reward-free RL” framework. In the exploration phase, the agent first collects trajectories from an MDP $\mathcal{M}$ without a pre-specified reward function. After exploration, it is tasked with computing near-optimal policies under for $\mathcal{M}$ for a collection of given reward functions. This framework is particularly suitable when there are many reward functions of interest, or when the reward function is shaped by an external agent to elicit desired behavior.
原文 arXiv:2002.02794;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2002.02794v1