Off-Policy Estimation of Long-Term Average Outcomes with Applications to Mobile Health
Peng Liao Department of Statistics, University of Michigan Predrag Klasnja School of Information, University of Michigan Susan Murphy Department of Statistics, Harvard University
Abstract
Due to the recent advancements in wearables and sensing technology, health scientists are increasingly developing mobile health (mHealth) interventions. In mHealth interventions, mobile devices are used to deliver treatment to individuals as they go about their daily lives. These treatments are generally designed to impact a near time, proximal outcome such as stress or physical activity. The mHealth intervention policies, often called just-in-time adaptive interventions, are decision rules that map a individual’s current state (e.g., individual’s past behaviors as well as current observations of time, location, social activity, stress and urges to smoke) to a particular treatment at each of many time points. The vast majority of current mHealth interventions deploy expert-derived policies. In this paper, we provide an approach for conducting inference about the performance of one or more such policies using historical data collected under a possibly different policy. Our measure of performance is the average of proximal outcomes over a long time period should the particular mHealth policy be followed. We provide an estimator as well as confidence intervals. This work is motivated
中文速览
移动健康干预(mHealth intervention)正越来越多地被用于慢性病管理,但这些干预策略大多由专家凭经验制定,其长期效果究竟如何,缺乏数据驱动的科学评估手段。这篇论文针对"离线策略评估"(off-policy evaluation)问题,提出了一套利用已有历史数据——这些数据是在某个行为策略(如随机试验中的随机化方案)下收集的——来推断另一套目标策略长期表现的统计方法;核心思路是将系统建模为马尔可夫决策过程(Markov Decision Process),以长期平均奖励(average reward)作为性能指标,通过估计值函数和平稳分布来构造点估计量并推导其渐近分布,从而给出置信区间。在真实数据应用中,作者用HeartSteps项目42天微随机试验(Micro-Randomized Trial)的数据,成功评估了多种活动建议推送策略在长达一年尺度上的预期效果。这项工作的重要意义在于,它为移动健康领域从"拍脑袋定策略"迈向"用数据选策略"提供了严格的统计推断工具,尤其适用于需要长期管理的慢性病场景。
原文 arXiv:1912.13088;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1912.13088v3