Comparing Dynamics: Deep Neural Networks versus Glassy Systems
Marco Baity-Jesi Levent Sagun Mario Geiger Stefano Spigler Gérard Ben Arous Chiara Cammarota Yann LeCun Matthieu Wyart Giulio Biroli
Abstract
We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are (1) the complexity of the loss landscape and of the dynamics within it, and (2) to what extent DNNs share similarities with glassy systems. Our findings, obtained for different architectures and datasets, suggest that during the training process the dynamics slows down because of an increasingly large number of flat directions. At large times, when the loss is approaching zero, the system diffuses at the bottom of the landscape. Despite some similarities with the dynamics of mean-field glassy systems, in particular, the absence of barrier crossing, we find distinctive dynamical behaviors in the two cases, showing that the statistical properties of the corresponding loss and energy landscapes are different. In contrast, when the network is under-parametrized we observe a typical glassy behavior, thus suggesting the existence of different phases depending on whether the network is under-parametrized or over-parametrized.
中文速览
训练深度神经网络(DNN)为何能绕开无数糟糕的局部极小值、顺利收敛,这一问题至今众说纷纭。研究者借用统计物理中研究"玻璃态系统"的工具——单点与双点时间关联函数——来追踪不同架构(从单隐层玩具模型到ResNet)在MNIST和CIFAR上的训练动力学。结果发现,训练过程中系统的减速并非源于翻越能量势垒,而是损失函数景观中平坦方向(flat directions)数量不断增加所致;训练后期,系统在损失景观底部附近做扩散运动,呈现出一种类似物理中"老化"(aging)的非平衡动力学现象。尽管过参数化网络与平均场玻璃模型在早期动力学上有相似之处,但二者在景观底部的行为存在本质差异,表明损失景观的几何结构与玻璃态能量景观并不相同;而当网络欠参数化时,则会出现真正的玻璃态行为,由此揭示出深度学习中过参数化与欠参数化两种截然不同的"相"。这一工作为理解神经网络为何在高维非凸优化中表现良好提供了来自统计物理的新视角。
原文 arXiv:1803.06969;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.06969v2