Benign Overfitting in Linear Regression
Peter L. Bartlett Department of Statistics, UC Berkeley, 367 Evans Hall, Berkeley CA 94720-3860 Computer Science Division, UC Berkeley, 387 Soda Hall, Berkeley CA 94720-1776 Philip M. Long Google Gábor Lugosi Economics and Business, Pompeu Fabra University; ICREA, Pg. Lluís Companys 23, 08010 Barcelona, Spain; Barcelona Graduate School of Economics Alexander Tsigler Department of Statistics, UC Berkeley, 367 Evans Hall, Berkeley CA 94720-3860
Abstract
The phenomenon of benign overfitting is one of the key mysteries uncovered by deep learning methodology: deep neural networks seem to predict well, even with a perfect fit to noisy training data. Motivated by this phenomenon, we consider when a perfect fit to training data in linear regression is compatible with accurate prediction. We give a characterization of linear regression problems for which the minimum norm interpolating prediction rule has near-optimal prediction accuracy. The characterization is in terms of two notions of the effective rank of the data covariance. It shows that overparameterization is essential for benign overfitting in this setting: the number of directions in parameter space that are unimportant for prediction must significantly exceed the sample size. By studying examples of data covariance properties that this characterization shows are required for benign overfitting, we find an important role for finite-dimensional data: the accuracy of the minimum norm interpolating prediction rule approaches the best possible accuracy for a much narrower range of properties of the data distribution when the data lies in an infinite dimensional space versus when th
中文速览
深度神经网络在完美拟合带噪声训练数据的同时仍能给出准确预测,这一"良性过拟合(benign overfitting)"现象令人困惑。为了从理论上解释它,本文聚焦最简单的情形——线性回归,研究最小范数插值估计量(minimum norm interpolating estimator,即用伪逆求得的零训练误差解)何时能保持接近最优的预测精度。核心发现是:良性过拟合发生的充要条件,可以用协方差矩阵(covariance matrix)特征值的两种"有效秩(effective rank)"来刻画——简单说,参数空间中"对预测不重要的低方差方向"的数量必须远超样本量,噪声才能被藏进这些方向而不破坏预测。进一步分析还揭示了一个有趣的维度效应:在无穷维空间中,良性过拟合只在极窄的特征值衰减速率范围内成立;而当数据处于维度随样本量增长的有限维空间时,容许良性过拟合的特征值衰减模式则宽泛得多。这一结果为理解过参数化模型为何能泛化提供了严格的有限样本理论基础,也暗示大规模有限维数据是深度学习中良性过拟合的典型场景。
原文 arXiv:1906.11300;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1906.11300v3