Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari1, Anna Choromanska2, Stefano Soatto1, Yann LeCun3,4, Carlo Baldassi5, Christian Borgs6, Jennifer Chayes6, Levent Sagun3, Riccardo Zecchina5 1 Computer Science Department, University of California, Los Angeles 2 Department of Electrical and Computer Engineering, New York University 3 Courant Institute of Mathematical Sciences, New York University 4 Facebook AI Research, New York 5 Dipartimento di Scienza Applicata e Tecnologia, Politecnico di Torino 6 Microsoft Research New England, Cambridge Email:
Abstract
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenvalues. We leverage upon this observation to construct a local-entropy-based objective function that favors well-generalizable solutions lying in large flat regions of the energy landscape, while avoiding poorly-generalizable solutions located in the sharp valleys. Conceptually, our algorithm resembles two nested loops of SGD where we use Langevin dynamics in the inner loop to compute the gradient of the local entropy before each update of the weights. We show that the new objective has a smoother energy landscape and show improved generalization over SGD using uniform stability, under certain assumptions. Our experiments on convolutional and recurrent networks demonstrate that Entropy-SGD compares favorably to state-of-the-art techniques in terms of generalization error and training time.
中文速览
深度神经网络训练时,模型找到的"好解"往往藏在损失曲面又宽又平的山谷里,而不是又窄又尖的最优点,宽谷中的解对参数扰动更鲁棒、泛化能力更强。为了让优化算法主动"偏爱"这类宽谷,作者提出了 Entropy-SGD 算法:核心思路是把原始损失替换为一个称作"局部熵(local entropy)"的目标函数,它同时衡量一个位置的损失深度和周围区域的平坦程度,从而在优化时自动绕开尖锐极小值、奔向宽阔平坦区域。算法实现上采用两层嵌套循环,内层用随机梯度朗之万动力学(stochastic gradient Langevin dynamics, SGLD)对局部熵的梯度进行高效近似,外层则按常规方式更新网络权重。在卷积网络和循环网络上的实验表明,Entropy-SGD 的泛化误差与当前最优方法相当甚至更好,同时在循环网络任务上训练速度最高可达 SGD 的两倍,为大规模深度学习提供了一条兼顾效果与效率的新路径。
原文 arXiv:1611.01838;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1611.01838v5