Scaling description of generalization with number of parameters in deep learning
Mario Geiger Institute of Physics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Arthur Jacot Institute of Mathematics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Stefano Spigler Institute of Physics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Franck Gabriel Institute of Mathematics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Levent Sagun Institute of Physics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Stéphane d’Ascoli Laboratoire de Physique Statistique, École Normale Supérieure, PSL Research University, 75005 Paris, France Giulio Biroli Laboratoire de Physique Statistique, École Normale Supérieure, PSL Research University, 75005 Paris, France Clément Hongler Institute of Mathematics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland Matthieu Wyart Institute of Physics, École Polytechnique Fédérale de Lausanne, 1015 Lausanne, Switzerland
Abstract
Supervised deep learning involves the training of neural networks with a large number $N$ of parameters. For large enough $N$ , in the so-called over-parametrized regime, one can essentially fit the training data points. Sparsity-based arguments would suggest that the generalization error increases as $N$ grows past a certain threshold $N^{*}$ . Instead, empirical studies have shown that in the over-parametrized regime, generalization error keeps decreasing with $N$ . We resolve this paradox through a new framework. We rely on the so-called Neural Tangent Kernel, which connects large neural nets to kernel methods, to show that the initialization causes finite-size random fluctuations $\|f_{N}-\bar{f}_{N}\|\sim N^{-1/4}$ of the neural net output function $f_{N}$ around its expectation $\bar{f}_{N}$ . These affect the generalization error $\epsilon_{N}$ for classification: under natural assumptions, it decays to a plateau value $\epsilon_{\infty}$ in a power-law fashion $\sim N^{-1/2}$ . This description breaks down at a so-called jamming transition $N=N^{*}$ . At this threshold, we argue that $\|f_{N}\|$ diverges. This result leads to a plausible explanation for the cusp in test err
中文速览
深度神经网络在参数数量远超训练数据点时(过参数化regime)仍能保持良好的泛化性能,这与传统统计学习理论的预测相悖——按照奥卡姆剃刀的稀疏性逻辑,参数越多应该越容易过拟合,但实验偏偏观察到测试误差随参数量增加而持续下降,只在"堵塞转变点"(jamming transition)N*附近出现一个误差峰值。本文借助神经切线核(Neural Tangent Kernel, NTK)这一将大型神经网络与核方法相连接的工具,从理论上证明:随机初始化会导致网络输出函数在其期望值附近产生幅度约为N^{-1/4}的随机涨落,这些涨落使分类测试误差以N^{-1/2}的幂律速度衰减至一个平台值,而在N*处涨落发散则自然解释了误差尖峰的成因。在MNIST和CIFAR图像数据集上的大量实验完整验证了上述预测,同时理论还给出一个实用建议:在固定计算资源的前提下,训练多个参数量略超N*的中等规模网络并对其输出取集成平均,比单纯堆大一个网络能获得更低的泛化误差。
原文 arXiv:1901.01608;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1901.01608v5