The Impact of Regularization on High-dimensional Logistic Regression
Fariborz Salehi Department of Electrical Engineering California Institute of Technology Pasadena, CA 91125 Ehsan Abbasi Department of Electrical Engineering California Institute of Technology Pasadena, CA 91125 Babak Hassibi Department of Electrical Engineering California Institute of Technology Pasadena, CA 91125
Abstract
Logistic regression is commonly used for modeling dichotomous outcomes. In the classical setting, where the number of observations is much larger than the number of parameters, properties of the maximum likelihood estimator in logistic regression are well understood. Recently, Sur and Candes [27] have studied logistic regression in the high-dimensional regime, where the number of observations and parameters are comparable, and show, among other things, that the maximum likelihood estimator is biased. In the high-dimensional regime the underlying parameter vector is often structured (sparse, block-sparse, finite-alphabet, etc.) and so in this paper we study regularized logistic regression (RLR), where a convex regularizer that encourages the desired structure is added to the negative of the log-likelihood function. An advantage of RLR is that it allows parameter recovery even for instances where the (unconstrained) maximum likelihood estimate does not exist. We provide a precise analysis of the performance of RLR via the solution of a system of six nonlinear equations, through which any performance metric of interest (mean, mean-squared error, probability of support recovery, etc.)
中文速览
高维数据中逻辑回归(logistic regression)的最大似然估计存在严重偏差,而现实场景里参数往往还具有稀疏等结构特性,已有研究既无法精确刻画这种偏差,也难以充分利用结构信息。本文引入凸正则化项来约束参数结构,提出正则化逻辑回归(Regularized Logistic Regression, RLR)框架,并借助凸高斯极大极小定理(Convex Gaussian Min-max Theorem, CGMT)将估计量的渐近行为精确归结为一组六元非线性方程组的解,从而可以显式计算均值、均方误差、支撑集恢复概率等任意局部Lipschitz性能指标,还能据此找到最优正则化参数。以$\ell_2^2$正则化和$\ell_1$稀疏正则化为典型案例,理论预测与大量数值仿真高度吻合。这是首个在样本数与参数维数可比的高维体制下精确表征正则化逻辑回归性能的工作,为实际应用中调优正则化参数提供了坚实的理论依据。
原文 arXiv:1906.03761;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1906.03761v4