The role of regularization in classification of high-dimensional noisy Gaussian mixture
Francesca Mignacco Université Paris-Saclay, CNRS, CEA, Institut de physique théorique, 91191, Gif-sur-Yvette, France Florent Krzakala Laboratoire de Physique de l’Ecole normale supérieure, ENS, Université PSL, CNRS, Sorbonne Université, Université de Paris, F-75005 Paris, France Yue M. Lu John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA 02138, USA Lenka Zdeborová Université Paris-Saclay, CNRS, CEA, Institut de physique théorique, 91191, Gif-sur-Yvette, France
Abstract
We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge, hinge and logistic regression, in the high-dimensional limit where the number $n$ of samples and their dimension $d$ go to infinity while their ratio is fixed to $\alpha=n/d$ . We discuss surprising effects of the regularization that in some cases allows to reach the Bayes-optimal performances. We also illustrate the interpolation peak at low regularization, and analyze the role of the respective sizes of the two clusters.
中文速览
在高维数据场景下,即便知道两个高斯簇(Gaussian clusters)的真实中心,分类器也不可避免地会犯一定比例的错误——这篇论文正是要精确刻画在这种"噪声不可消除"的情况下,岭回归、铰链损失、逻辑回归等经典正则化凸分类器的泛化误差究竟是多少。作者利用 Gordon 极小极大不等式,在样本数 $n$ 与维度 $d$ 同比例趋于无穷($\alpha = n/d$ 固定)的极限下,严格推导出了闭合形式的渐近公式,覆盖任意凸损失函数、任意正则化强度以及两个簇大小不相等的情形。研究发现了几个反直觉的现象:适当的正则化有时能让经验风险最小化达到贝叶斯最优(Bayes-optimal)性能,而弱正则化区域则会出现泛化误差先升后降的"插值峰值"(interpolation peak / double descent);此外,一个类似 Hebb 规则的简单加权平均估计器,在调好偏置后可以直接达到贝叶斯最优。这些结果在一个统一的理论框架内严格解释了当前机器学习界热议的多个高维统计现象,为理解正则化的作用机制提供了坚实的数学基础。
原文 arXiv:2002.11544;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2002.11544v1