Double/Debiased Machine Learning for Treatment and Structural Parameters
Victor Chernozhukov† Denis Chetverikov‡ Mert Demirer† Esther Duflo† Christian Hansen§ Whitney Newey† James Robins⋆ †Massachusetts Institute of Technology, 50 Memorial Drive, Cambridge, MA, 02139, USA †University of California Los Angeles, 315 Portola Plaza, Los Angeles, CA 90095 §University of Chicago, 5807 S. Woodlawn Ave., Chicago, IL 60637 ⋆ Harvard University, 677 Huntington Avenue Boston, Massachusetts 02115
Abstract
We revisit the classic semiparametric problem of inference on a low dimensional parameter $\theta_{0}$ in the presence of high-dimensional nuisance parameters $\eta_{0}$ . We depart from the classical setting by allowing for $\eta_{0}$ to be so high-dimensional that the traditional assumptions, such as Donsker properties, that limit complexity of the parameter space for this object break down. To estimate $\eta_{0}$ , we consider the use of statistical or machine learning (ML) methods which are particularly well-suited to estimation in modern, very high-dimensional cases. ML methods perform well by employing regularization to reduce variance and trading off regularization bias with overfitting in practice. However, both regularization bias and overfitting in estimating $\eta_{0}$ cause a heavy bias in estimators of $\theta_{0}$ that are obtained by naively plugging ML estimators of $\eta_{0}$ into estimating equations for $\theta_{0}$ . This bias results in the naive estimator failing to be $N^{-1/2}$ consistent, where $N$ is the sample size. We show that the impact of regularization bias and overfitting on estimation of the parameter of interest $\theta_{0}$ can be removed by usin
中文速览
当我们想在高维数据中用机器学习方法(如随机森林、Lasso、神经网络)估计某个低维因果参数(比如处理效应)时,直接把机器学习得到的辅助函数估计值代入估计方程,会因为正则化偏差和过拟合引入严重偏误,导致估计量的收敛速度远低于统计推断所要求的根号N速度。本文提出"双重/去偏机器学习"(Double/Debiased ML,DML)方法,核心是两个简单却关键的操作:一是使用对辅助参数不敏感的Neyman正交矩条件来估计目标参数,从而消除正则化偏差;二是采用交叉拟合(cross-fitting)这一高效数据分割方式,防止过拟合造成的偏误渗入目标参数的估计。理论证明,DML估计量能以根号N速度收敛、近似无偏且渐近正态,从而支持有效的置信区间构造,且对所用机器学习方法几乎没有限制。这一框架统一覆盖了部分线性回归、工具变量、平均处理效应等多种常见因果推断场景,为在大数据和高维控制变量环境下做可靠因果推断提供了坚实的理论基础和实用工具。
原文 arXiv:1608.00060;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1608.00060v7