Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes
Nathan Kallus Department of Operations Research and Information Engineering and Cornell Tech Cornell University New York, NY 10044, USA Masatoshi Uehara Department of Statistics Harvard University Cambridge, MA 02138, USA
Abstract
Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time the semiparametric efficiency limits of OPE in Markov decision processes (MDPs), where actions, rewards, and states are memoryless. We show existing OPE estimators may fail to be efficient in this setting. We develop a new estimator based on cross-fold estimation of $q$ -functions and marginalized density ratios, which we term double reinforcement learning (DRL). We show that DRL is efficient when both components are estimated at fourth-root rates and is also doubly robust when only one component is consistent. We investigate these properties empirically and demonstrate the performance benefits due to harnessing memorylessness.
中文速览
用离线历史数据来评估一个新策略的好坏(即离线策略评估,OPE),是强化学习里一个核心却棘手的问题,难点在于如何在不增加额外探索成本的前提下尽可能精确地估计目标策略的累积回报。本文首次从半参数效率(semiparametric efficiency)角度切入,系统研究了马尔可夫决策过程(MDP)与非马尔可夫决策过程(NMDP)这两种模型下OPE的最优估计精度下界,发现现有估计量普遍未能利用马尔可夫"无记忆性",导致效率损失,在长时间步下甚至呈指数级恶化(即"时间步灾难")。为此,作者提出了"双重强化学习"(Double Reinforcement Learning,DRL)估计量,通过交叉折叠方式联合学习 q 函数和边际密度比,将二者插入高效影响函数来构造估计;理论上证明了只要两个成分各自以四次方根速率收敛,DRL 就能达到半参数效率下界,且即便其中一个成分估计不一致,估计量仍保持一致性(双重稳健性)。这一结果意味着在真实的MDP场景中,可以用灵活的机器学习方法估计中间量,同时获得统计最优的策略评估精度,对医疗、教育等高风险决策场景具有重要的实践价值。
原文 arXiv:1908.08526;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1908.08526v3