Deep Reinforcement Learning at the Edge of the Statistical Precipice
Rishabh Agarwal Thanks: Outstanding Paper Award. Correspondence to Rishabh Affiliation: Google Research, Brain Team Affiliation: MILA, Université de Montréal Max Schwarzer Affiliation: MILA, Université de Montréal Pablo Samuel Castro Affiliation: Google Research, Brain Team Aaron Courville Affiliation: MILA, Université de Montréal Marc G. Bellemare Affiliation: Google Research, Brain Team
Abstract
Deep reinforcement learning (RL) algorithms are predominantly evaluated by comparing their relative performance on a large suite of tasks. Most published results on deep RL benchmarks compare point estimates of aggregate performance such as mean and median scores across tasks, ignoring the statistical uncertainty implied by the use of a finite number of training runs. Beginning with the Arcade Learning Environment (ALE), the shift towards computationally-demanding benchmarks has led to the practice of evaluating only a small number of runs per task, exacerbating the statistical uncertainty in point estimates. In this paper, we argue that reliable evaluation in the few-run deep RL regime cannot ignore the uncertainty in results without running the risk of slowing down progress in the field. We illustrate this point using a case study on the Atari 100k benchmark, where we find substantial discrepancies between conclusions drawn from point estimates alone versus a more thorough statistical analysis. With the aim of increasing the field’s confidence in reported results with a handful of runs, we advocate for reporting interval estimates of aggregate performance and propose performance
原文 arXiv:2108.13264;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2108.13264v4