The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari Institute for Computational and Mathematical Engineering, Stanford UniversityDepartment of Electrical Engineering and Department of Statistics, Stanford University
Abstract
Deep learning methods operate in regimes that defy the traditional statistical mindset. Neural network architectures often contain more parameters than training samples, and are so rich that they can interpolate the observed labels, even if the latter are replaced by pure noise. Despite their huge complexity, the same architectures achieve small generalization error on real data.
中文速览
用随机特征(random features)回归模型来严格刻画深度学习中"双下降"(double descent)现象,一直缺乏既足够真实又能精确求解的理论工具。作者在高维球面上用 N 个随机神经元做岭回归(ridge regression),并在 N、n、d 同比例趋于无穷的极限下,精确推导出测试误差的解析表达式。结果表明:测试误差随模型复杂度先呈 U 形上升、在插值阈值处达到峰值,随后再次下降,且当信噪比超过某临界值时,参数量远大于样本量的极度过参数化插值器反而取得全局最优泛化误差,零正则化(ridgeless)方案有时也优于任何有限正则化。这是首个无需人为设置模型误设结构、就能完整再现双下降全部特征的可解析计算模型,为理解过参数化神经网络为何能良好泛化提供了坚实的数学基础。
原文 arXiv:1908.05355;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1908.05355v5