Gradient Descent Finds Global Minima of Deep Neural Networks
Simon S. Du Affiliation: Machine Learning Department, Carnegie Mellon University Correspondence to: Jason D. Lee Affiliation: Data Science and Operations Department, University of Southern California Haochuan Li Affiliation: School of Physics, Peking University Affiliation: Center for Data Science, Peking University, Beijing Institute of Big Data Research Liwei Wang Affiliation: Key Laboratory of Machine Perception, MOE, School of EECS, Peking University Affiliation: Center for Data Science, Peking University, Beijing Institute of Big Data Research Xiyu Zhai Affiliation: Department of EECS, Massachusetts Institute of Technology
Abstract
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero training loss in polynomial time for a deep over-parameterized neural network with residual connections (ResNet). Our analysis relies on the particular structure of the Gram matrix induced by the neural network architecture. This structure allows us to show the Gram matrix is stable throughout the training process and this stability implies the global optimality of the gradient descent algorithm. We further extend our analysis to deep residual convolutional neural networks and obtain a similar convergence result.
原文 arXiv:1811.03804;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1811.03804v4