A Modern Maximum-Likelihood Theory for High-dimensional Logistic Regression
Pragya Sur Department of Statistics, Stanford University, Stanford, CA 94305, U.S.A. Emmanuel J. Candès Department of Mathematics, Stanford University, Stanford, CA 94305, U.S.A.
Abstract
Every student in statistics or data science learns early on that when the sample size $n$ largely exceeds the number $p$ of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there are formulas to predict the variability of these estimates which are used for the purpose of statistical inference; for instance, to produce p-values for testing the significance of regression coefficients. Although these formulas come from large sample asymptotics, we are often told that we are on reasonably safe grounds when $n$ is large in such a way that $n\geq 5p$ or $n\geq 10p$ . This paper shows that this is far from the case, and consequently, inferences routinely produced by common software packages are often unreliable.
原文 arXiv:1803.06964;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.06964v4