Regularization in High-Dimensional Regression and Classification via Random Matrix Theory
Panagiotis Lolas
Abstract
We study general singular value shrinkage estimators in high-dimensional regression and classification, when the number of features and the sample size both grow proportionally to infinity. We allow models with general covariance matrices that include a large class of data generating distributions. As far as the implications of our results are concerned, we find exact asymptotic formulas for both the training and test errors in regression models fitted by gradient descent, which provides theoretical insights for early stopping as a regularization method. In addition, we propose a numerical method based on the empirical spectra of covariance matrices for the optimal eigenvalue shrinkage classifier in linear discriminant analysis. Finally, we derive optimal estimators for the dense mean vectors of high-dimensional distributions. Throughout our analysis we rely on recent advances in random matrix theory and develop further results of independent mathematical interest.
中文速览
高维数据中,样本量和特征维度相当时,经典统计理论会严重失效,如何为更广泛的"奇异值压缩估计量"给出精确的渐近误差公式,是本文要解决的核心问题。作者借助随机矩阵理论(random matrix theory),在特征维度与样本量之比趋于固定常数的高维渐近框架下,推导出了涵盖一大类谱压缩函数的回归预测误差精确极限,以及线性判别分析(linear discriminant analysis)分类误差的精确渐近公式。基于这些结果,论文给出了梯度下降训练过程中训练误差与测试误差随迭代步数演化的显式表达式,从理论上解释了早停(early stopping)作为正则化手段的有效性;同时提出了用于最优特征值压缩分类器和高维均值估计的数值方法。这项工作不仅将岭回归(ridge regression)的最优性结论推广到了更一般的压缩估计量,还为深度学习中"过训练"现象提供了严格的数学支撑,对理解高维统计推断的根本机制具有重要意义。
原文 arXiv:2003.13723;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2003.13723v1