Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation
Ziyang Tang University of Texas at Austin The first two authors contributed equally to this work. Yihao Feng* University of Texas at Austin Lihong Li Google Research Dengyong Zhou Google Research Qiang Liu University of Texas at Austin
Abstract
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationary density ratio, but at the cost of introducing potentially high biases due to the error in density ratio estimation. In this paper, we develop a bias-reduced augmentation of their method, which can take advantage of a learned value function to obtain higher accuracy. Our method is doubly robust in that the bias vanishes when either the density ratio or the value function estimation is perfect. In general, when either of them is accurate, the bias can also be reduced. Both theoretical and empirical results show that our method yields significant advantages over previous methods.
中文速览
无限时域(infinite-horizon)的离线策略评估(off-policy evaluation)长期受制于重要性采样权重随时间步指数级膨胀带来的高方差问题,Liu等人(2018a)通过直接估计目标策略与行为策略之间的平稳状态密度比来绕开这一问题,但密度比估计的误差又会引入较大偏差。本文在此基础上引入值函数估计作为辅助,构造了一个双重鲁棒(doubly robust)估计器:只要密度比或值函数中任意一个估计准确,整体偏差就会消失;即便两者都不完美,只要其中一个较准,偏差也能得到有效压缩。理论分析揭示了该方法与密度函数和值函数之间原始-对偶关系的深层联系,实验结果也证实其在多种任务上的精度显著优于已有方法,为实际应用中可靠的离线策略评估提供了更坚实的工具。
原文 arXiv:1910.07186;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1910.07186v1