A Modern Maximum-Likelihood Theory for High-dimensional Logistic Regression
Pragya Sur Department of Statistics, Stanford University, Stanford, CA 94305, U.S.A. Emmanuel J. Candès Department of Mathematics, Stanford University, Stanford, CA 94305, U.S.A.
Abstract
Every student in statistics or data science learns early on that when the sample size $n$ largely exceeds the number $p$ of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there are formulas to predict the variability of these estimates which are used for the purpose of statistical inference; for instance, to produce p-values for testing the significance of regression coefficients. Although these formulas come from large sample asymptotics, we are often told that we are on reasonably safe grounds when $n$ is large in such a way that $n\geq 5p$ or $n\geq 10p$ . This paper shows that this is far from the case, and consequently, inferences routinely produced by common software packages are often unreliable.
中文速览
逻辑回归(logistic regression)在统计学教材中有一条金科玉律:只要样本量 n 够大(哪怕只是变量数 p 的5到10倍),最大似然估计(MLE)就近似无偏,教科书给的标准误公式也基本靠谱。这篇论文用理论与模拟双管齐下证明这条金科玉律其实大错特错:当 n 和 p 同比例增长时,MLE 会系统性地把效应量夸大,其真实波动远超经典公式的预测,而被软件普遍用来生成 p 值的似然比检验(LRT)也根本不服从卡方分布——这意味着日常科研中大量依赖统计软件输出结果的推断其实是不可靠的。作者由此建立了一套新的渐近理论,核心发现是:上述三种失真现象的修正量仅依赖于一个标量——信号强度 γ,而非整个高维参数向量,因而只需从数据中估计这一个参数就能大幅校正推断结果。这项工作为现代高维数据分析中逻辑回归的可信使用提供了理论基础和可操作的修正方案,对医学、社会科学等大量依赖逻辑回归的领域具有直接的现实意义。
原文 arXiv:1803.06964;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.06964v4